One thing that teaching hospitals have is a longer ladder of experience. For example many hospitals have medical students, interns, residents, chief residents, fellows and attendings. One idea we contemplated was to have "treatment" start on the cheapest possible model and see how it ended up, only escalating to more capable/expensive models as necessary. We rejected this idea because we suspected that it would end up in more time/tokens on the capable models to fix any problems.
In the end, medical students learn to become doctors. Cheaper models never learn to become more capable models, so the metaphor is not apt.
Did you do any ablation studies on what is actually useful vs what happens to just work because these systems can work around whatever people do?
I compare this to OpenAI's symphony prompt which, at a high level, does exactly the same thing outlined here except model selection and doesn't really have the need for roles or hospital metaphors.
No, we haven't performed any ablation studies yet - it's a good suggestion.
The comparison with OpenAI's Symphony is valid (the two systems do broadly the same thing). One thing that's different about Sinai is that it keeps the code author and reviewer in separate agents with separate context. At the time we built it, we suspected that this would lead to better outcomes, but again, we haven't validated that it does.
I confess that there's a lot more validation we could do with this model but we just haven't found the time. One thing we don't cover in the post is that this is a side project for both of us, so we have less time than we'd like to perform experimentation and validation.
In the end, medical students learn to become doctors. Cheaper models never learn to become more capable models, so the metaphor is not apt.
I compare this to OpenAI's symphony prompt which, at a high level, does exactly the same thing outlined here except model selection and doesn't really have the need for roles or hospital metaphors.
The comparison with OpenAI's Symphony is valid (the two systems do broadly the same thing). One thing that's different about Sinai is that it keeps the code author and reviewer in separate agents with separate context. At the time we built it, we suspected that this would lead to better outcomes, but again, we haven't validated that it does.
I confess that there's a lot more validation we could do with this model but we just haven't found the time. One thing we don't cover in the post is that this is a side project for both of us, so we have less time than we'd like to perform experimentation and validation.