AI NewsProductivityAnnouncement
Cockroach Labs ran coding agents 5 months, 1M lines, 7 reverts
Cockroach Labs ran a coding agent pipeline modelled on a teaching hospital for five months, merging 1,238 pull requests across more than a million lines of code, with seven reverts and about $135,000 spent on Claude tokens.

Image: Cockroach Labs
Why it mattersA team on a database reports what rework rate, token spend and revert count their coding agents produced over five months, which gives any other team real numbers to compare its own pipeline against.
A database company publishing five months of its own coding-agent numbers tells you more about the state of agent-written code than any vendor benchmark.
On Thursday, Cockroach Labs wrote up MOLT Sinai, a coding agent pipeline modelled on a teaching hospital that it has been running against its MOLT migration tools since April 21. The post, written by Jacob Lacouture and Zach Lite, reports that the pipeline merged 1,238 pull requests across more than one million lines of code over about five months, with seven reverts, and consumed around $135,000 in Claude tokens over that period.
The teaching-hospital pipeline
Each agent plays a named role: a Triage Nurse, a Fellow who writes the diagnostic workup and the treatment plan, a Review Attending who looks for faults in both, a Discharge Nurse who audits whether the review was done, and a human Chief who takes escalations. Four more agents run outside the main flow: a Charge Nurse that revives stalled issues every 30 minutes, an Infection Control agent that locks the pipeline when main breaks, a Safety Department that writes weekly process reviews, and a Research Department that proposes new work. Everything runs on GitHub Actions, with labels on the issue tracking which stage each patient is in.
A few rules carry most of the safety. A Fellow may not write code before a reviewed plan. Any workup that exceeds 1,000 lines of code must be decomposed into sub-issues. An agent may not weaken a test to make it pass. When stuck, a Fellow writes an I-PASS handoff, a format borrowed from clinical shift changes, and the receiving agent must say what it understood before starting work. Approved human decisions go into an append-only precedent log that other agents cite instead of re-escalating.
What the numbers show
A single large experiment opened the pipeline on 21 April: adding IBM Db2 support to MOLT. A planning agent split the parent issue into 15 sub-issues, two of which were decomposed again, and by close the pipeline had opened 32 sub-issues under it, merged 27 pull requests, sent work back 55 times and escalated nine issues, two to a human. The post says the equivalent Oracle work in 2024 took nine months and about $160,000 in engineering time; the Db2 token bill was $4,172.
Across the full five months, Cockroach Labs says nearly half of all issues were filed by the hospital for itself. Each issue cost about $84 on average and spent one to two days in the pipeline, often waiting on a human reviewer.
The post also records limits that mattered. One urgent one-line fix took 11 rework rounds over two days through false claims in the pull-request description, test-quality objections and two rebases, after which the team added a circuit breaker that hands off to an Attending after four consecutive new-finding rounds or six rounds of any kind. The Discharge Nurse was found to load about 21,800 words of instructions, roughly 29,000 tokens, on every run before reading the pull request, and a July audit of the 25 skill files flagged 23 percent of their 100,000 words as removable without changing a gate.
A team deciding whether to adopt a similar structure now has two numbers to work with: a measured method for keeping agent-written code acceptable on a correctness-critical codebase, and a measured bill for doing so.
Source
- Cockroach Labs, Jacob Lacouture and Zach Lite, Five months treating bugs like patients and coding agents like a medical team
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.


