Relic turns multi-agent coding failures into persistent executable protocols

The paper starts from a common failure in multi-agent software work. One coding agent changes an interface in a repository, while another keeps developing against the old version, and existing tests go stale. A conversation can settle that one episode. But when the participants change, what makes the lesson keep governing the team?
The authors' answer is Relic, which turns recurring collaboration failures into organization-owned, executable protocols. Members reflect on visible work, propose rules, and govern their adoption. Once adopted, a protocol binds triggers, responsibilities, required evidence and execution consequences to the runtime. It stays open to revision and retirement. In one traced case, repeated integration friction produced an interface-review rule that governed later pull requests and was revised as work continued.
The evidence comes in three parts. First, across 360 controlled runs over ten software workloads and three models, Relic raised complete-contract delivery from 14.06% to 19.76% (stated in the paper as +5.71 percentage points) over a matched structured team without the protocol lifecycle. It improved all four verified production endpoints in every model stratum.
Second, the paper tests fresh-member transfer, where new members join a team. Behavioral correctness was 25.4% with no inherited protocol, 34.6% when the same rules were provided as readable text, and 41.2% with executable bindings. The paper describes that as a +6.5-point advantage over text alone.
Third, on the full CooperBench benchmark, after excluding broken benchmark pairs, Relic achieves 367/477 (76.9%). The authors say this is the best reported result among peer-structured systems. On the fixed 48-pair same-model subset, Relic also exceeds Solo (29/48 vs. 26/48), reversing the coordination loss exhibited by the official peer baseline.
The authors conclude that these results show how collaboration experience can become persistent organizational state that remains useful beyond the members who created it.
Key facts
- Relic turns recurring collaboration failures among agents into organization-owned, executable protocols that members propose, adopt, revise and retire.
- Across 360 controlled runs (ten software workloads, three models), complete-contract delivery rose from 14.06% to 19.76% versus a matched structured team without the protocol lifecycle.
- Under fresh-member transfer, behavioral correctness was 25.4% with no inherited protocol, 34.6% with rules as readable text, and 41.2% with executable bindings.
- On CooperBench (broken pairs excluded), Relic scores 367/477 (76.9%), which the authors call the best reported result among peer-structured systems.
- On the fixed 48-pair same-model subset, Relic beats Solo, 29/48 vs. 26/48.
Why it matters
Teams of coding agents can collide, for example when one changes an interface and another keeps building on the old one. Fixing a single conflict in conversation does not help once the participants change. Relic's idea is to move the lesson out of the conversation and into a protocol owned by the organization, which the runtime enforces and which can still be revised or retired. The paper frames this as collaboration experience becoming persistent organizational state.
Who it affects
The setting is multi-agent software teams, where several coding agents work in the same repository and hand work to each other. Researchers and builders of such systems are the natural audience, especially those who care about what happens when team members are replaced. The paper's tests cover ten software workloads and the CooperBench benchmark.
How to use it
The abstract describes a workflow rather than a product. Members reflect on visible work, propose rules, and govern their adoption. Adopted rules bind triggers, responsibilities, required evidence and execution consequences to the runtime. The reported fresh-member results suggest that executable bindings work better than handing new members the same rules as text. No code or dataset release is mentioned.
How solid is it
The evidence is broad for an abstract: 360 controlled runs across ten workloads and three models, a fresh-member transfer comparison, and results on CooperBench. The gain in complete-contract delivery is modest in absolute terms, from 14.06% to 19.76%. On the 48-pair subset the margin over Solo is 29 against 26. The three models are not named, and the ten workloads are not described beyond their number. It is not stated whether the results are peer-reviewed. The stated differences (+5.71 points and +6.5 points) are as printed in the paper and do not exactly equal the differences of the listed rates. The claim of a best reported result applies to peer-structured systems on the full CooperBench after excluding broken pairs.
Risks and caveats
Everything here rests on the authors' own abstract. The figures for fresh-member transfer and CooperBench come without stated trial counts. Absolute rates stay low in the main comparison: even with Relic, complete-contract delivery is 19.76%. The traced interface-review case is a single example. No timescale or cost of running Relic is given.
“Adopted protocols bind triggers, responsibilities, required evidence, and execution consequences to the runtime, while remaining open to revision and retirement.”
— Relic paper abstract