Sharpening Tax paper: RL post-training costs agents solution coverage

Sharpening Tax paper: RL post-training costs agents solution coverage

A paper on the Hugging Face papers page (2610.01509) takes on a hypothesis about reinforcement learning (RL) post-training of large language models: that it merely sharpens behaviors the base model already has, improving single-shot accuracy at the cost of solution coverage. That trade-off has been observed in math and coding tasks. The authors say it need not carry over to agentic tasks, where multi-turn tool use and interaction may require capabilities newly acquired during post-training.

Their test points the other way. The authors call it a surprising finding: pre-trained LLMs, equipped with a light inference harness, can serve as capable agents. Their single-shot accuracy (pass@1) is far lower than that of post-trained counterparts. Yet given a sufficient test-time budget, the base models often surpass the post-trained ones in solution coverage (pass@K).

The paper then analyzes the mechanism. According to the authors, post-training pushes tasks toward two extremes, always solved or never solved. That improves sampling efficiency and consistency, but it costs solution coverage.

To measure that cost, the authors propose Sharpening Tax, a diagnostic metric that quantifies the loss in test-time scalability after post-training. They evaluated it across 14 base/post-trained model pairs from four families and three agentic benchmarks, 42 cases in total. The tax is prevalent in most settings, can be estimated from a few rollouts, and correlates well with other metrics.

Finally, the paper presents posterior-tempered group sampling (PTGS), a simple plug-and-play Bayesian sampler that adapts the sampling temperature per prompt to its estimated difficulty. Applied during RL training in two agentic environments, PTGS paid a smaller tax than the fixed-temperature baseline. It solved more tasks under repeated sampling and also improved single-shot accuracy.

Key facts

  • The authors report that base LLMs with a light inference harness often beat their post-trained counterparts on pass@K (solution coverage) in agentic tasks, despite far lower pass@1.
  • Mechanism: post-training pushes tasks toward two extremes, always solved or never solved, which improves sampling efficiency and consistency but cuts coverage.
  • Sharpening Tax is a diagnostic metric for the loss in test-time scalability after post-training; it was measured across 14 base/post-trained pairs from four families and three agentic benchmarks (42 cases).
  • The tax is prevalent in most settings, can be estimated from a few rollouts, and correlates well with other metrics.
  • PTGS, a Bayesian sampler that adapts temperature per prompt to estimated difficulty, paid a smaller tax than a fixed-temperature baseline in RL training in two agentic environments.

Why it matters

The common view is that RL post-training only sharpens what a base model can already do, trading coverage for single-shot accuracy. That had been seen in math and coding. This paper tests whether it also holds for agentic tasks, and reports that it does: base models, given enough samples, often solve more tasks than post-trained ones. The paper also gives a way to put a number on the loss, which is hard to discuss without a measure.

Who it affects

Teams that post-train LLMs with RL for agentic work are the direct audience, since the paper describes a cost they may be paying without measuring it. People who deploy agents with repeated sampling at test time are affected too, because the result bears on whether a post-trained model or its base model scales better with a larger sampling budget.

How to use it

The paper offers two practical pieces. Sharpening Tax can be estimated from a few rollouts, so checking a base/post-trained pair does not need a huge run. PTGS is described as a simple plug-and-play Bayesian sampler that sets the sampling temperature per prompt according to its estimated difficulty, and it was applied during RL training. The source does not mention code or release availability.

How solid is it

The evidence base is broad for a single paper: 14 base/post-trained model pairs from four families, three agentic benchmarks, 42 cases in total. The claim is hedged by the authors themselves: the tax is prevalent in most settings, not all. The PTGS result comes from two agentic environments. The abstract gives no numeric values for pass@1, pass@K, the size of the tax or PTGS gains, and does not name the model families or benchmarks, so the size of the effects cannot be judged from it.

Risks and caveats

The base-model advantage appears at pass@K only when the test-time budget is sufficient, and pass@1 is far lower, so it does not help where only one attempt is allowed. The size of the needed budget (the value of K) is not stated. The PTGS evidence covers only two agentic environments, and the findings are the authors' own claims from the abstract.

“Our surprising finding is that pre-trained LLMs, equipped with a light inference harness, can serve as capable agents.”

— Paper abstract, Hugging Face papers 2610.01509