One thing I've been experimenting with in my own agent orchestration system is an agent that hangs around and does post-merge acceptance testing after the work ships. Any plans to add a follow-up phase? Travel nurse?
It does! It helps getting promotion with tokenmaxxing.
Once we started to review the code in more detail (and put the hospital into "human approval mode" to stop the momentum) we found one major thing that we corrected.
When issues were decomposed into a DAG, sibling issues would often defer work that they thought that their sibling(s) would handle, and that work sometimes just got dropped. This was because there was no way for sibling issues to communicate with each other. We've since added a deferred scope ledger and policy. If any issue is going to defer scope it must comment about the deferral in the code, and write it to a special ledger with the rationale, and how/when the deferral should be actioned. This has prevented several items from falling through the cracks, but we're constantly refining what is a valid deferral and what should be fixed immediately.
There are several other things that we've corrected over time, which makes me think that another blog post is in order. The full list would be too extensive to try and address in the comments section.
Could you structure the DAG so that after each node that contains the work, you have one dependent node that verifies each distinct requirement was implemented as expected?
This way if a single requirement is dropped the system alerts you rather than it being silently dropped.
It makes sense intuitively that if a task has nothing that depends on it the LLM might accidentally attempt to drop it (even purposefully as an optimization).
It was definitely human driven, but I believe that the agents did the actual trimming.
That type of encoding (starting with the statement: you're at a teaching hospital) is profoundly more efficient than having to deeply and reliably articulate what each of the various roles your agents embody are, let alone their interactions.
Perhaps there is more room for any of us to consider what existing systems or identities are described abundantly in training sets and can be leveraged to ritualistically encode these social/civic/cultural dynamics saying: "act as though you're ___." Not exactly a new point, but one this write-up certainly is pulling me towards.
Thank you for sharing and the care you put into writing this!
I will confess however that some people internally find the model confusing. For example, one user couldn't remember that to get an issue actioned, they needed to put it into the "waiting room". They wanted to be able to action their issues without learning how complex medical systems operate. There seems to be more work on our plates to make this resonate with all users.
i kinda want to try to build this off of github. it's essentially just an event bus / message queue that workers (agents) tap into.
probably what will work best in the future is specific training or fine-tuning against curated datasets of narrative fiction and/or texts in general? not really sure.
Fable feels like overkill for this also.
I use Redmine (lightly customized) and have found that to work very well. Especially for performing review / UAT of the work or providing response to the agents questions when it wants me to pick an option.
The underlying data model of a ticket / project management system is so rich and well suited to bodies of work.
AI Software Factory: $4172 / 2 days = $2086/day
This seems unsustainable, unless you're also generating 3x the revenue.
In any case the way to think of this is not "/day".
The reason is simple. If you buy 1000 barrels of oil, you buy 1000 barrels of oil not 17.4 days of oil. There is now a disconnect between work done and time. Infact you would be sane if you said "that result I can get in 2 days for $4172, if you can get me that same result in 1 hour, I'd pay $8344". See where this is going?
Yes a lot of thought work is now a commodity, and if you want the commodity faster (last minute booking, uber to come quicker etc.) you pay more not less. Value being $/hour is over.
[1] The ? acts as both a question and a regex.
Modern best practices for a medical team (standard roles, responsibilities, consequences, processes, etc.) have evolved through trial and error, since the advent of civilization, to prevent costly human error.
In hindsight, it's not too surprising that the same best practices can be applied by teams of AI agents to minimize costly AI error.
It's still incredible. We sure live in interesting times!
In the end, medical students learn to become doctors. Cheaper models never learn to become more capable models, so the metaphor is not apt.
I compare this to OpenAI's symphony prompt which, at a high level, does exactly the same thing outlined here except model selection and doesn't really have the need for roles or hospital metaphors.
The comparison with OpenAI's Symphony is valid (the two systems do broadly the same thing). One thing that's different about Sinai is that it keeps the code author and reviewer in separate agents with separate context. At the time we built it, we suspected that this would lead to better outcomes, but again, we haven't validated that it does.
I confess that there's a lot more validation we could do with this model but we just haven't found the time. One thing we don't cover in the post is that this is a side project for both of us, so we have less time than we'd like to perform experimentation and validation.