Cockroach Labs' agent hospital built Db2 support for MOLT in under two days

Cockroach Labs' agent hospital built Db2 support for MOLT in under two days

Cockroach Labs describes an internal system called MOLT Sinai, a pipeline of coding agents built to look like a teaching hospital. Its headline result: between a Wednesday evening and a Friday afternoon in April, MOLT, the tool the company ships for migrating databases to CockroachDB, gained the ability to migrate from IBM Db2 sources. The company says no human wrote any of the code, and the token bill was $4172.

The scope was large. Db2 has a rich SQL dialect, a complex type system and a native wire protocol, so support meant a new schema converter, a new Fetch path for pulling rows out and loading them into CockroachDB, a new Verify path for comparing them afterwards, a vendored ANTLR grammar, a Docker image for CI, and more than ten thousand lines of test fixtures. When the company added Oracle support to MOLT in 2024, the equivalent work took 9 months and cost around $160K in engineering time. The source puts the Db2 run at 164x faster and 38x cheaper.

The run began with a single GitHub issue describing the ask. A planning agent judged it too large to tackle at once and decomposed it into fifteen sub-issues with an explicit dependency graph: foundation, type system, row iterator, Fetch, Verify, Convert, CI, then test data. Two sub-issues were judged too large in their own workups and were split again. The agents also filed another dozen issues against their own work, including fixture gaps, a type-mapping bug and an isolation-level fix. By the time the parent issue closed, thirty-two sub-issues had been opened under it, twenty-seven pull requests had merged, reviewers had sent work back fifty-five times, and nine issues had been escalated, two of them to a human. The authors say test coverage, measured against PostgreSQL (their best-tested dialect), was equivalent.

The design came first. The company chose a hospital model because its priority is correctness, not throughput: a subtle bug in a migration tool could corrupt a customer's data on its way into a database meant to maintain correctness above all else. The authors contrast it with throughput-optimized efforts such as Gas Town, which runs many agents in parallel with a merge-handling layer, and a Cursor experiment that rebuilt SQLite so fast it needed its own version control system. The question they cared about was how many merged PRs they would have been embarrassed to have merged themselves. They say outright that the hospital is not cheaper than a swarm or simpler than a single agent, but that it produces a higher-quality product than either.

The vocabulary is medical: issues are patients, merging is discharge, and the humans in charge are the Chiefs of Medicine. Everything runs on GitHub Actions. A patient's state lives in labels on its issue plus scratch files the hospital keeps. Each stage is a workflow triggered by a 'label in'; when the agent finishes it writes a 'label out', which triggers the next stage. Each agent loads a role-specific prompt and leaves a structured note in the issue's chart, marked with a sentinel such as SINAI:TREATMENT_PLAN so later agents can find it. Around the core flow sit a Charge Nurse who rounds every thirty minutes looking for stuck patients, an Infection Control agent that can lock down every workflow if main breaks, a Safety Department that writes weekly process reviews and runs post-incident conferences, and a Research Department that proposes new work.

A handful of rules encoded in the agents' skills provide most of the safety. No code before a plan, no plan without review: the Fellow must reproduce the problem, perform a differential diagnosis, run targeted investigation and post a treatment plan (diagnosis, files to change, tests to add, risks, open questions), which a Review Attending, a different agent instance running a prompt built around finding fault, approves or rejects. Don't improvise: if scope changes, the Fellow stops, posts a revised plan and returns to plan review. No shortcuts on testing: an agent may not disable, skip, weaken or modify a test to make it pass, and reviews start by disabling the change and checking that the new tests fail. Never flail: a stuck Fellow writes an I-PASS handoff (Illness severity, Patient summary, Action list, Situation awareness, Synthesis) and escalates, and the receiver must write back what it understood. The Discharge Nurse audits the reviewer: the last gate before merge does not re-review the code, it verifies that an approval exists, the review template was filled out, no threads are unresolved, CI is green and the commit history is clean. Finally, agents must check an append-only precedent log before escalating to the Chief; only a human can create precedents, and over five months this has cut how often agents escalate.

The company also reports results from a mirror of the MOLT repo, covering April 21 to September 11. In just under five months the pipeline landed well over a million lines of code, split into changes small enough for agents to review confidently. Nearly half the issues filed in the repo were created by the hospital for itself: decomposition children, follow-ups, Research Department proposals and Safety Department corrective actions. A human decides what gets admitted, but humans stopped being the main source of work in the second month.

Db2 was the first large feature through the hospital, and the authors call it an experiment. After the first week of autonomous operation they turned on a 'human approval mode' that requires a human to look at every merge, and say only in that mode could they be certain the code was good enough to ship to customers. In that mode the hospital built what it was created for: the Migration Assistant, an AI-based tool that walks users through a complete Postgres-to-CockroachDB migration, including schema conversion, data load, verification and routine conversion. It is in preview for Postgres migrations and was almost entirely written by the hospital's agents.

On cost, MOLT Sinai has consumed just over $135,000 in Claude tokens since April, which works out to around $84 for the average patient (a GitHub issue). Issues typically spend one to two days in the hospital, often dominated by waiting on a human reviewer.

Key facts

  • Cockroach Labs says IBM Db2 source support was added to MOLT in less than two days, with no human-written code, for a $4172 token bill; the equivalent Oracle work in 2024 took 9 months and about $160K in engineering time.
  • The pipeline, MOLT Sinai, runs entirely on GitHub Actions with hospital roles (Fellow, Review Attending, Discharge Nurse, Charge Nurse, Infection Control) and rules such as no code before a reviewed plan and no weakening tests.
  • The Db2 parent issue produced thirty-two sub-issues and twenty-seven merged pull requests; reviewers sent work back fifty-five times and nine issues were escalated, two of them to a human.
  • Over April 21 to September 11 the pipeline landed well over a million lines of code in a mirror repo, at just over $135,000 in Claude tokens, or around $84 per average issue.
  • The authors say the design favours quality over throughput and is not cheaper than a swarm or simpler than a single agent; they added human approval mode after the first week of autonomous operation.

Why it matters

Most public experiments in agent-written code have optimized for speed. Cockroach Labs argues the real problem is trusting agents to produce production-grade, well-tested, in-scope code, and it answers with process rather than volume: mandatory plans, mandatory second opinions and an audit of the reviewer. The headline comparison is stark. Db2 support took less than two days and a $4172 token bill, against 9 months and around $160K in engineering time for Oracle in 2024, which the source puts at 164x faster and 38x cheaper. The authors also report that the pipeline generated nearly half of its own backlog and that humans stopped being the main source of work in the second month.

Who it affects

Teams building automated coding pipelines, especially where a bug is expensive, are the obvious audience: the article is a worked example of a quality-first design borrowed from high-stakes medicine, including I-PASS handoffs. For Cockroach Labs customers, MOLT now supports IBM Db2 as a migration source, and the Migration Assistant for Postgres migrations is in preview and was almost entirely written by the hospital's agents. Engineers who review agent output are affected too: in this setup humans decide what gets admitted, make precedent-setting calls as Chiefs of Medicine, and often become the bottleneck, since issues spend one to two days in the hospital, frequently waiting on a human reviewer.

How to use it

The article reads as a design to copy rather than a product to install. The mechanics it gives: model each stage as a GitHub workflow triggered by a label, keep patient state in issue labels and scratch files, load a role-specific prompt per stage, and have agents write structured notes with a sentinel marker so later agents can parse them. Encode the rules in skills: no code before a plan, no plan without review, no test changes to make a test pass, a handoff format for stuck agents, a final gate that checks the review rather than redoing it, and a precedent log agents must consult before escalating. The authors also recommend a human approval mode before shipping to customers; they turned theirs on after the first week of autonomous operation. No pricing or licence for any tooling is given.

How solid is it

This is a first-party case study from the company that built and ships the tool, so the numbers are its own and no outside party has checked them. The figures are specific and internally consistent: $4172 against about $160K is roughly 38x, and the source states both multiples. Two caveats on the comparison. The $4172 is the token bill only; it does not include the cost of building the hospital or human time spent on it. The $160K Oracle figure is stated as engineering time, not token cost, and the basis of the estimate is not given. The claim that test coverage was equivalent is measured against PostgreSQL, the company's best-tested dialect. The Db2 two-day run is not stated to have gone through human approval mode; the text says that mode was turned on after the first week of autonomous operation.

Risks and caveats

The authors concede the hospital brings bureaucracy, wait times and cost, which they judge worthwhile for database work. The total spend is large in absolute terms: just over $135,000 in Claude tokens since April, about $84 for the average issue. Autonomy had limits: even in the Db2 run, nine issues were escalated and two went to a human, and the authors say they could only be certain of shippable quality once every merge had human eyes on it. They also note that agents sometimes do not do what they are told, which is why the Discharge Nurse exists. Self-generated work is a double-edged number: nearly half of the issues in the mirror repo were created by the hospital for itself. No defect rates, bug counts or post-merge incident figures are given in the visible text, and the year of the Db2 work is not stated.

“We'd much rather have the agents take longer to run (even 10x longer) than land the wrong data in a customer migration to CockroachDB.”

— Cockroach Labs, MOLT Sinai write-up