Opera, a verbal critic for coding agents, lifts resolve rates by up to 15 points

Opera, a verbal critic for coding agents, lifts resolve rates by up to 15 points

A paper on Hugging Face presents Opera, a verbal critic framework for long-horizon coding agents. The starting point is a problem with feedback: coding agents need timely corrections, but feedback can be ineffective or even harmful when it misjudges ongoing work or fails to address the underlying problem. The authors say existing critics focus on evaluating trajectories and generating feedback, but rarely track what happens after the feedback is delivered.

Opera's answer is to treat each correction as a persistent note that is followed until the diagnosed problem is resolved. It decides when to review through periodic and event-driven triggers, diagnoses issues with typed operators, audits its feedback against visible evidence before delivering it, and tracks the agent's subsequent actions to tell mere compliance from actual resolution.

As a test-time critic, Opera improves the resolve rate of non-critic agents by up to 12.4 percentage points on Terminal-Bench 2.1, up to 15.0 on a SWE-Bench Pro subset, and up to 8.9 on DeepSWE v1.1, across four policy models. The authors also report that it achieves the highest mean resolve rate among competitive critic baselines on all three benchmarks, and that it improves policy models when the policy critiques itself.

The framework is also used beyond inference. Opera-guided rollouts provide approximately on-policy training data. Fine-tuning Qwen3.5-9B on them improves its resolve rate on held-out SWE-Bench Pro repositories by 10.2 percentage points without a critic at inference time. The authors say this matches fine-tuning on rollouts from a stronger model, and that the fine-tuned model preserves its performance when the harness is switched from Openhands to Terminus-2, which the abstract describes as substantially degrading. The code is available at the GitHub repository dongyuanjushi/Opera.

Key facts

  • Opera is a verbal critic framework for long-horizon coding agents that treats each correction as a persistent note, followed until the diagnosed problem is resolved.
  • It reviews on periodic and event-driven triggers, diagnoses issues with typed operators, audits feedback against visible evidence, and tracks whether the agent actually resolved the problem or only complied.
  • As a test-time critic it improves non-critic agents' resolve rates by up to 12.4 points on Terminal-Bench 2.1, up to 15.0 on a SWE-Bench Pro subset and up to 8.9 on DeepSWE v1.1, across four policy models.
  • Fine-tuning Qwen3.5-9B on Opera-guided rollouts improves its resolve rate on held-out SWE-Bench Pro repositories by 10.2 points with no critic at inference time.
  • The authors report the highest mean resolve rate among competitive critic baselines on all three benchmarks; code is on GitHub at dongyuanjushi/Opera.

Why it matters

Long-horizon coding agents work for many steps, and a critic that nudges them mid-run can help or hurt. The paper's point is that most critics stop at generating feedback and never check what the agent did with it. Opera adds that missing step: it keeps a correction open until the diagnosed problem is resolved, and separates an agent that merely followed the advice from one that fixed the issue. It is also used as a source of training data, so the same mechanism serves both inference and fine-tuning.

Who it affects

Teams building or evaluating long-horizon coding agents, and researchers working on critics, agent feedback and training data for software engineering tasks. The reported results cover Terminal-Bench 2.1, a SWE-Bench Pro subset and DeepSWE v1.1, and the fine-tuning experiment uses a 9B-parameter open model, Qwen3.5-9B.

How to use it

The authors released their code at https://github.com/dongyuanjushi/Opera. The paper describes two uses: Opera as a test-time critic that guides an agent during a run, and Opera-guided rollouts as approximately on-policy training data for fine-tuning a policy model, after which no critic is needed at inference time. No cost, latency or compute overhead of running Opera is stated.

How solid is it

These are the authors' own results, reported in a paper abstract with code released for others to inspect. The 12.4, 15.0 and 8.9 point gains are maximums ('up to') across four policy models, not averages or typical gains. The abstract names no authors or institutions, and gives only percentage-point gains, not absolute resolve rates. It claims the best mean resolve rate among competitive critic baselines on all three benchmarks.

Risks and caveats

The headline gains are best cases, so typical improvement may be smaller. The four policy models are not named, except Qwen3.5-9B in the fine-tuning experiment, and the stronger model used as the comparison for fine-tuning is not named either. The abstract does not say how large the harness-switch degradation is, nor numerically how much of the fine-tuned model's performance is preserved. The 10.2 point fine-tuning gain is stated only for held-out SWE-Bench Pro repositories, with no baseline given. Overhead of running the critic is not stated.

“Existing critics focus on evaluating trajectories and generating feedback, but rarely track what happens after feedback is delivered.”

— Opera paper abstract