Hacker News

Favorites Setup
Five months treating bugs like patients and coding agents like a medical team (cockroachlabs.com)
2026-10-09 Fri | 145 points by rafiss | original
[−]ajstorm · 2026-10-09 Fri 15:29 UTC · link
Rafi and I, who authored this post, will be hanging out here for any questions people may have.
[−]contingencies · 2026-10-10 Sat 23:11 UTC · link
Which inherent limitations did you recognize in the metaphor before commencing this research?
[−]ajstorm · 2026-10-11 Sun 01:15 UTC · link
One thing that teaching hospitals have is a longer ladder of experience. For example many hospitals have medical students, interns, residents, chief residents, fellows and attendings. One idea we contemplated was to have "treatment" start on the cheapest possible model and see how it ended up, only escalating to more capable/expensive models as necessary. We rejected this idea because we suspected that it would end up in more time/tokens on the capable models to fix any problems.

In the end, medical students learn to become doctors. Cheaper models never learn to become more capable models, so the metaphor is not apt.

[−]Eridrus · 2026-10-11 Sun 00:00 UTC · link
Did you do any ablation studies on what is actually useful vs what happens to just work because these systems can work around whatever people do?

I compare this to OpenAI's symphony prompt which, at a high level, does exactly the same thing outlined here except model selection and doesn't really have the need for roles or hospital metaphors.

[−]ajstorm · 2026-10-11 Sun 01:08 UTC · link
No, we haven't performed any ablation studies yet - it's a good suggestion.

The comparison with OpenAI's Symphony is valid (the two systems do broadly the same thing). One thing that's different about Sinai is that it keeps the code author and reviewer in separate agents with separate context. At the time we built it, we suspected that this would lead to better outcomes, but again, we haven't validated that it does.

I confess that there's a lot more validation we could do with this model but we just haven't found the time. One thing we don't cover in the post is that this is a side project for both of us, so we have less time than we'd like to perform experimentation and validation.

[−]ArtRichards · 2026-10-11 Sun 08:39 UTC · link
Yes agreed, using a system very similar without the hospital vibe :)
[−]what · 2026-10-11 Sun 03:21 UTC · link
The blog post links to an issue in the Sinai repo, but it’s private or just doesn’t exist?
[−]losteric · 2026-10-11 Sun 03:33 UTC · link
are any of these artifacts public and available for inspection/use?
[−]ajstorm · 2026-10-11 Sun 04:15 UTC · link
Not currently, but open sourcing this has been discussed. Stay tuned.
[−]jordanlewis · 2026-10-09 Fri 16:44 UTC · link
Great post!

One thing I've been experimenting with in my own agent orchestration system is an agent that hangs around and does post-merge acceptance testing after the work ships. Any plans to add a follow-up phase? Travel nurse?

[−]ajstorm · 2026-10-09 Fri 16:49 UTC · link
That's a very interesting idea. Sounds more like a "routine follow-up" in the medical model.
[−]Spooky23 · 2026-10-10 Sat 22:47 UTC · link
Reminds me of the “surgical team” development model in the Mythical Man Month.
[−]sroerick · 2026-10-11 Sun 00:52 UTC · link
I thought this too. It's fun that they were working with DB2 as well.
[−]james_marks · 2026-10-10 Sat 22:58 UTC · link
A hint at why GitHub actions have been unreliable. How many teams are running factories like this on GH infra now?
[−]gchamonlive · 2026-10-10 Sat 23:52 UTC · link
This is Microsoft scale, I'd be surprised if it made any difference these labs. It's more likely it's plain managerial mishandling of the infra in chasing new profit heights.
[−]devin · 2026-10-11 Sun 00:11 UTC · link
I posted in another threat about this: I am seeing a lot of people building their own little bespoke factories. They introduce endless quality gates until it slows development down, and then they add more agents to decide when to run certain actions, and on and on. The end result from what I've seen and personally participated in, is that it often winds up providing negative value in the software development lifecycle. It creates a whole lot of heat, but IMO is not helping the teams utilizing them to ship value any faster than they would with a more limited setup.
[−]supermdguy · 2026-10-11 Sun 02:15 UTC · link
I've gone through this cycle recently. I think static lint/type checks are super useful, but agentic code review loops can easily go off the rails.
[−]zx8080 · 2026-10-11 Sun 04:08 UTC · link
> not helping

It does! It helps getting promotion with tokenmaxxing.

[−]devin · 2026-10-11 Sun 04:57 UTC · link
lol, no doubt.
[−]Zanfa · 2026-10-11 Sun 06:29 UTC · link
I just finished ripping out one of these “dark factory” setups. Removed about 70k lines of code and 750k words of generated documentation. For what is effectively a 5-screen app.
[−]git_rancher · 2026-10-10 Sat 23:48 UTC · link
The patient “leaves” when the bug is fixed?
[−]tough · 2026-10-11 Sun 00:02 UTC · link
What would be the analogy if the patient dies?
[−]unrented7977 · 2026-10-11 Sun 00:11 UTC · link
Bug becomes a feature
[−]kinduff · 2026-10-11 Sun 00:13 UTC · link
Since we are all building our version of this, what mistakes did you made until you arrived to a well-balanced solution? I really like that you know how much an issue is, because then you can start optimizing.
[−]ajstorm · 2026-10-11 Sun 00:57 UTC · link
I'd say that the first "mistake" we made was having it run in auto-merge mode. It was incredible to see what it could produce, and the speed with which it worked, but while the results seemed good, they were being produced at a rate that we couldn't human-verify. This is not to say that they were bad, but that we had no way to convince ourselves that they were good.

Once we started to review the code in more detail (and put the hospital into "human approval mode" to stop the momentum) we found one major thing that we corrected.

When issues were decomposed into a DAG, sibling issues would often defer work that they thought that their sibling(s) would handle, and that work sometimes just got dropped. This was because there was no way for sibling issues to communicate with each other. We've since added a deferred scope ledger and policy. If any issue is going to defer scope it must comment about the deferral in the code, and write it to a special ledger with the rationale, and how/when the deferral should be actioned. This has prevented several items from falling through the cracks, but we're constantly refining what is a valid deferral and what should be fixed immediately.

There are several other things that we've corrected over time, which makes me think that another blog post is in order. The full list would be too extensive to try and address in the comments section.

[−]aetherspawn · 2026-10-11 Sun 02:19 UTC · link
Is a code comment and ledger the best way? Should the agent just fill out a form or something and attach it to the sub issue. This is how the hospital would work.
[−]ajstorm · 2026-10-11 Sun 03:36 UTC · link
It often does that too, but we keep the ledger and the code comment as well, to ensure that if it doesn't get resolved by the sibling, that it's not lost. If the sibling does resolve it, it's removed from the ledger and the code.
[−]dkubb · 2026-10-11 Sun 06:40 UTC · link
> When issues were decomposed into a DAG, sibling issues would often defer work that they thought that their sibling(s) would handle, and that work sometimes just got dropped.

Could you structure the DAG so that after each node that contains the work, you have one dependent node that verifies each distinct requirement was implemented as expected?

This way if a single requirement is dropped the system alerts you rather than it being silently dropped.

It makes sense intuitively that if a task has nothing that depends on it the LLM might accidentally attempt to drop it (even purposefully as an optimization).

[−]Veelox · 2026-10-11 Sun 00:17 UTC · link
You give a very precise measure of redundancy in the skills. Can you give a bit more detail in how you decided you needed to audit them and how you went about it? Was it fully agent driven? Mostly human?
[−]ajstorm · 2026-10-11 Sun 00:44 UTC · link
All that credit goes to Rafi. I believe that he either noticed that they were getting long winded in a code review, or suspected that they needed trimming after reading this blog post from Anthropic: https://claude.dev/blog/the-new-rules-of-context-engineering....

It was definitely human driven, but I believe that the agents did the actual trimming.

[−]Veelox · 2026-10-11 Sun 00:58 UTC · link
That is helpful. Thank you :)
[−]dingaling911 · 2026-10-11 Sun 00:21 UTC · link
Maybe I missed it, but I didn't really see anything about the long term quality or maintainability of the code. All I see is agent agent agent.
[−]ajstorm · 2026-10-11 Sun 00:49 UTC · link
It's true that we don't have long-term maintainability data just yet. We've just shipped the first product of this model to customers and likely won't have any detailed maintainability data for several months (and for good data, several years). We hope to author more blog posts on this experiment in the future.
[−]Kinrany · 2026-10-11 Sun 03:03 UTC · link
Perhaps doing a random sample of the steps by hand will be a good way to notice maintainability issues?
[−]finnborge · 2026-10-11 Sun 00:22 UTC · link
This feels to me like a lovely example of both high-quality systems design and information theory. From my perspective, one of the more abstractly interesting thoughts surfaced by this write-up is the degree to which you've "encoded" a highly complex set of relationships through use of metaphor.

That type of encoding (starting with the statement: you're at a teaching hospital) is profoundly more efficient than having to deeply and reliably articulate what each of the various roles your agents embody are, let alone their interactions.

Perhaps there is more room for any of us to consider what existing systems or identities are described abundantly in training sets and can be leveraged to ritualistically encode these social/civic/cultural dynamics saying: "act as though you're ___." Not exactly a new point, but one this write-up certainly is pulling me towards.

Thank you for sharing and the care you put into writing this!

[−]ajstorm · 2026-10-11 Sun 00:40 UTC · link
Thanks for the thoughtful response. We too were surprised by how deeply we could take the model of the teaching hospital to software engineering. Whenever we think to expand the system in one dimension or another, the teaching hospital model seems to have a nearby analogy.

I will confess however that some people internally find the model confusing. For example, one user couldn't remember that to get an issue actioned, they needed to put it into the "waiting room". They wanted to be able to action their issues without learning how complex medical systems operate. There seems to be more work on our plates to make this resonate with all users.

[−]jacquesm · 2026-10-11 Sun 01:44 UTC · link
Now imagine what it is like to run an actual hospital with real people whose lives (or the lives of their loved ones) are often at stake, with underpaid and overworked staff and with messy biological creatures as the subjects instead of bits and bytes. If anything this whole exercise should also give you a much deeper appreciation of the people that feel themselves called to help others.
[−]jaggederest · 2026-10-11 Sun 02:08 UTC · link
I had a similarly useful analogy of the legal system emerge, in a similar way. I think there's a great analogy between policy in the legal system and policy in software engineering, and they have some awfully good (and very, very historical) ways to think about e.g. amending, repealing, and adjudicating things based on those policies.
[−]slopinthebag · 2026-10-11 Sun 02:43 UTC · link
yeah this is really cool. i've been thinking about sort of a mirrored idea, where you end up with wizards, mages, clerics etc and fantasy terminology used. idk if it would be as effective as this is, but maybe it would be more creative somehow?

i kinda want to try to build this off of github. it's essentially just an event bus / message queue that workers (agents) tap into.

[−]baddash · 2026-10-11 Sun 03:05 UTC · link
i think role-play and utilizing the full power of language, stories, character, and narrative will unlock very sophisticated use-cases and in general a new dimension to agentic systems much in the way you're describing.

probably what will work best in the future is specific training or fine-tuning against curated datasets of narrative fiction and/or texts in general? not really sure.

[−]exe34 · 2026-10-11 Sun 05:52 UTC · link
I feel like this is something we do with humans as well. Things like "scrum", "sprint", "the clean coder", etc.
[−]zmj · 2026-10-11 Sun 01:00 UTC · link
Nice writeup. Structured handoffs and external plan reviews are good takeaways.
[−]ajstorm · 2026-10-11 Sun 01:15 UTC · link
Thanks! And thanks for reading.
[−]mncharity · 2026-10-11 Sun 02:57 UTC · link
One role I didn't see was patient advocate/representative? That might be another approach to non-convergence - "how is this going?" and escalation.
[−]ajstorm · 2026-10-11 Sun 03:27 UTC · link
We actually have a /sinai-advocate skill, where a human can advocate on behalf of a stuck patient. We use it every once in a while when the labels get screwed up, or a workflow fails for some reason.
[−]anonymous908213 · 2026-10-11 Sun 04:17 UTC · link
I can think of nothing I'd like to use less than a database or filesystem vibecoded by Gas Town-flavored psychosis. Roleplaying with LLMs is not the secret to producing amazing code.
[−]zshrdlu · 2026-10-11 Sun 06:03 UTC · link
Seems to me it's just research and experimentation.
[−]drc500free · 2026-10-11 Sun 04:19 UTC · link
I absolutely love how you are able to pull so much latent behavior from the underlying LLM. I wonder what other analogies can be pulled into agentic coding that come baked into the existing weights.
[−]fathermarz · 2026-10-11 Sun 05:14 UTC · link
I’m surprised at the reasoning behind using the most expensive model which likely led to the “doing too much”. I think it would have been nice to have A/B tested different hospital setups in order to maximize not only efficiency but to see where the strength of each model started to emerge in a given role.

Fable feels like overkill for this also.

[−]singularity2001 · 2026-10-11 Sun 05:27 UTC · link
congenially my agents started calling bugs gaps
[−]K0balt · 2026-10-11 Sun 05:37 UTC · link
I set different kinds of structures for different projects, Complete with setting and ambiance. It’s like agents work better if they are role-playing. It’s extremely disorienting and people with marginal mental stability are going to really have a bad time. What have we wrought?
[−]mimischi · 2026-10-11 Sun 06:35 UTC · link
If I wanted to build something like this, at least conceptually with the roles, where’d I start? My first guess would be to give Claude your blog post; but any other pointers to make it work reliably? Do you happen to have the system open source?
[−]reachableceo · 2026-10-11 Sun 09:45 UTC · link
I am curious why so many of these systems are based on GitHub issues. Why not use a proper ticket system?

I use Redmine (lightly customized) and have found that to work very well. Especially for performing review / UAT of the work or providing response to the agents questions when it wants me to pick an option.

The underlying data model of a ticket / project management system is so rich and well suited to bodies of work.

[−]davidmurdoch · 2026-10-11 Sun 10:47 UTC · link
Who defines what proper is?
[−]tujux · 2026-10-11 Sun 10:01 UTC · link
Humans: $160k / 9 months = $600/day

AI Software Factory: $4172 / 2 days = $2086/day

This seems unsustainable, unless you're also generating 3x the revenue.

[−]epolanski · 2026-10-11 Sun 10:04 UTC · link
Only one of the two trends is downwards.
[−]paid_dot_expert · 2026-10-11 Sun 10:18 UTC · link
I don't get your maths? [1] 160/9 months != 600 for any given number of days a week e.g. 5,6,7 (something between 6 and 7), but where did the 9 come from, is it some kind of adjustment for weekends? Anyway there is also a concept of "fully loaded employee cost".

In any case the way to think of this is not "/day".

The reason is simple. If you buy 1000 barrels of oil, you buy 1000 barrels of oil not 17.4 days of oil. There is now a disconnect between work done and time. Infact you would be sane if you said "that result I can get in 2 days for $4172, if you can get me that same result in 1 hour, I'd pay $8344". See where this is going?

Yes a lot of thought work is now a commodity, and if you want the commodity faster (last minute booking, uber to come quicker etc.) you pay more not less. Value being $/hour is over.

[1] The ? acts as both a question and a regex.

[−]cs702 · 2026-10-11 Sun 11:23 UTC · link
Great post. Thank you for sharing it on HN.

Modern best practices for a medical team (standard roles, responsibilities, consequences, processes, etc.) have evolved through trial and error, since the advent of civilization, to prevent costly human error.

In hindsight, it's not too surprising that the same best practices can be applied by teams of AI agents to minimize costly AI error.

It's still incredible. We sure live in interesting times!