VeriHarness turns an LLM into its own verifier for long-horizon agent tasks

VeriHarness turns an LLM into its own verifier for long-horizon agent tasks

As LLM agents take on longer and more complex tasks, checking their outputs gets harder. A paper on Hugging Face papers asks how verification can be strengthened with a fixed base model, without reference answers or grading rubrics at test time.

The starting point is repeated sampling. Running an agent several times produces multiple rollouts, and these can contain complementary correct claims. The difficulty is knowing which claims to trust. The authors first report an observation: disagreement among rollouts often exposes correct alternatives, while consensus can conceal errors.

That finding motivates VeriHarness. It turns the underlying LLM that a generator uses into an agentic verifier by giving it a workspace, evidence tools and reusable verification skills. Two roles do the checking. A disagreement resolver tests competing claims against environmental evidence. A consensus challenger tests the claims all rollouts share and searches for requirements the rollouts omitted. Their findings then guide the selection and revision of the final artifact.

The evaluation covers five long-horizon workspace benchmarks and two frontier models. The authors state that VeriHarness achieves the highest selection scores among the evaluated baselines. Evidence-backed revision improves average performance further: the gain over a single rollout reaches 6.2 points with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8. The authors also show that the verification skills can self-improve from failure feedback.

Finally, the authors release the full pool of approximately 26,000 rollouts across all five benchmarks and both models. Producing it cost over $100,000, and it is offered to support future research on agentic verification.

Key facts

  • VeriHarness makes the generator's own underlying LLM act as an agentic verifier, with a workspace, evidence tools and reusable verification skills.
  • A disagreement resolver checks competing claims against environmental evidence; a consensus challenger tests shared claims and looks for omitted requirements.
  • Across five long-horizon workspace benchmarks and two frontier models, the authors say it has the highest selection scores among the evaluated baselines.
  • Evidence-backed revision lifts average performance over a single rollout by 6.2 points with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8.
  • About 26,000 rollouts, produced at a cost of over $100,000, are released for future research.

Why it matters

Long-horizon agent tasks produce outputs that are hard to check, and in many settings there is no reference answer or rubric to grade against. This work targets exactly that gap, with a fixed base model and no answer key at test time. Its central finding is that disagreement among sampled rollouts often exposes correct alternatives, while agreement can hide errors. That cuts against the habit of treating the majority answer as the safe one.

Who it affects

Researchers and engineers building or evaluating LLM agents for long, multi-step workspace tasks are the direct audience. The paper evaluates Gemini 3.5 Flash and Claude Opus 4.8, so teams using models of that class are the ones the results speak to most directly. Researchers working on verification also get the released rollout pool.

How to use it

The abstract describes the pattern rather than a product: sample several rollouts, let the same LLM verify them with a workspace, evidence tools and verification skills, then select and revise the final artifact. The released pool of approximately 26,000 rollouts is meant for future research on agentic verification. No release link, license, or date for the rollout pool is given.

How solid is it

The claims come from the authors' own abstract, and the benchmarks, baselines and authors are not named there. The 6.2 and 6.4 point figures are gains in average performance from evidence-backed revision over a single rollout. They are not the margin by which VeriHarness beats other selection methods; the size of that margin is not given. No absolute scores are given, and the unit of the points (percent or other scale) is not stated.

Risks and caveats

The abstract gives no inference cost or overhead of running VeriHarness itself. The over $100,000 figure is the cost of producing the rollout pool, not the cost of the method per task. It also does not say how many rollouts are sampled per task, and results are not broken down by which of the two models a given selection result applies to. The claim that verification skills can self-improve from failure feedback is stated by the authors without further detail in the abstract.