GapFT fine-tunes on the Pass@K-Pass@1 gap to lift single-sample accuracy

An arXiv paper introduces GapFT, a fine-tuning method built around a simple observation about how language models are post-trained with verified answers. When an exact verifier is available, a model can be made to look better at test time by sampling several responses and picking one that passes. Other deployments use beam search, adaptive sampling or tools. The paper asks a narrower question: can that search-exposed behavior be absorbed into the model itself, so that single-sample decoding, where each query gets one response and no search, gets better?

The authors argue that existing verified-response post-training recipes do not generally distinguish problems the model already solves on its first decode from failures that it only recovers within K samples. Under a fixed budget, this can spend training examples repeating behavior the deployed policy already has.

GapFT changes which examples are used. It picks training evidence by the source checkpoint's single-sample outcome and fine-tunes on the Pass@K-Pass@1 gap: problems the policy fails on one sample but solves within K samples. The comparison is kept tight. The authors match training examples, processed tokens and optimizer steps, and keep the objective unchanged. GapFT fills the matched budget with recovered failures, and an exact decomposition separates corrections of recovered and missed failures from regressions on first-decode successes.

Results are reported on two logic benchmarks with Llama-3.1-8B. On LogiQA 2.0 and ReClor, GapFT improves Pass@1 by 14.4 and 13.9 points over the source model. It outperforms budget-matched uniform verified RFT at the same learning rate, and it matches fine-tuning on the full verified pool while using one third of the data. A single decode of the GapFT model matches the source model's verifier-selected Pass@4 accuracy.

Two further pieces of evidence are offered. A randomized control attributes the gains to covering distinct failures, and the authors' analysis relates available gains to transferable failure support. A three-seed replication with Qwen2.5-7B retains positive gains over uniform RFT on both logic tasks.

Key facts

  • GapFT fine-tunes only on the Pass@K-Pass@1 gap: problems the model fails on one sample but solves within K samples.
  • With Llama-3.1-8B, Pass@1 rises by 14.4 points on LogiQA 2.0 and 13.9 points on ReClor over the source model.
  • It matches fine-tuning on the full verified pool using one third of the data, and beats budget-matched uniform verified RFT at the same learning rate.
  • A single decode of the GapFT model matches the source model's verifier-selected Pass@4 accuracy.
  • A randomized control attributes the gains to covering distinct failures; a three-seed Qwen2.5-7B replication keeps positive gains over uniform RFT on both tasks.

Why it matters

Search at inference time (sampling many answers and letting a verifier pick one) costs compute on every query. The paper tries to move that benefit into the weights, so one decode does the work. The reported result is that a single GapFT decode matches the source model's verifier-selected Pass@4 accuracy. The idea behind it is about data selection: under a fixed budget, training on problems the model already solves on its first try repeats behavior it already has, so GapFT spends the budget on recovered failures instead.

Who it affects

Teams that post-train models on verified responses, especially where an exact verifier exists and where the deployed model answers with a single sample and no search. The paper's setting is narrow: Llama-3.1-8B and Qwen2.5-7B on two logic tasks, LogiQA 2.0 and ReClor.

How to use it

The recipe, as described, is to take the source checkpoint, sample K responses per problem, keep the problems it fails on one sample but solves within K, and fine-tune on those while holding examples, processed tokens, optimizer steps and the objective fixed against the baseline. The value of K used for training is not stated. No code or model release is mentioned.

How solid is it

The claims come from the authors' abstract. The comparison is designed to be fair: examples, tokens and optimizer steps are matched and the objective is unchanged. A randomized control supports the explanation that gains come from covering distinct failures, and a three-seed Qwen2.5-7B replication keeps positive gains over uniform RFT on both logic tasks. The abstract gives point improvements over the source model but not the absolute Pass@1 scores, and the size of the gain over uniform verified RFT is not quantified. The size of the Qwen2.5-7B gains is likewise not given, only that they are positive.

Risks and caveats

Only two logic tasks (LogiQA 2.0 and ReClor) and two models (Llama-3.1-8B, Qwen2.5-7B) are mentioned, with no results on other domains such as math or code. The authors themselves tie available gains to transferable failure support, which suggests the benefit depends on the failures being ones that training can carry over. The method also needs a source checkpoint's outcomes on the training problems, and an exact verifier to tell solved from unsolved.

“Under a fixed budget, this can spend examples repeating behavior the deployed policy already has.”

— From the paper's abstract