Keyword-matching benchmarks give small models false credit for tool use, paper finds

The paper argues that keyword-matching benchmarks can credit small models for tool use they never perform. To show it, the authors document a false positive in a matched-architecture pair of Spanish security language models. One has 661.6M parameters (called the 600M): about 65% of its data is code and technical text, and it had no dedicated SFT. The other has 1,109M parameters (called the 1B): it was trained on a web-heavy, multi-phase curriculum and then given 6B tokens of tool-SFT. The two share a decoder, tokenizer and special tokens, and on lenient tool-use metrics they score almost identically (B4: 0.660 for the 600M vs. 0.650 for the 1B).
Stricter checks pull them apart. In verbatim-reproduction checks on training examples, the 600M emits valid tool calls with generalized arguments on 6 of 6 examples, while the 1B does so on 0 of 6 across checkpoints. A first-token probe then localizes the 1B's failure to a missing prior: the probability of the first token <|tool_call|> is between 10^-4 and 10^-5. The authors say this prior was erased by the 1B's web-heavy training phase.
The authors then repair the 1B with a targeted SFT recipe: a diverse corpus, a learning rate 5x higher, 2,202 steps and about 3.3 GPU-hours. It uses three orders of magnitude fewer tokens than the failed phase. On all 269 corpus rows, valid emission rises from 0.100 to 0.959, against 0.926 for the 600M. On 238 unseen prompts, the repaired 1B passes 0.536 vs. 0.428 for the 600M (p = 0.004). Embedding-drift checks show the trigger token's tied embedding did not move: 97.7% of the bf16 embedding table remains bit-identical. The authors conclude the changes live in the surrounding network.
There are caveats. Both models over-trigger and rarely answer negative prompts without a call: the rate is 0.09 for the 600M and 0.17 for the repaired 1B. Factorial analyses confirm that all repair configurations install the format. Whether suppression benefits from a diverse corpus is still a hypothesis, because results are sensitive to the random seed. The authors say their diagnostic ladder costs minutes of CPU time and should gate tool-use claims on small models.
Key facts
- A 661.6M model and a 1,109M model with the same decoder, tokenizer and special tokens score 0.660 vs. 0.650 on the lenient tool-use metric B4.
- In verbatim-reproduction checks on training examples, the 600M emits valid tool calls on 6/6 examples and the 1B on 0/6 across checkpoints.
- A first-token probe puts the 1B's probability on <|tool_call|> at 10^-4 to 10^-5, a missing prior the authors attribute to its web-heavy training phase.
- A targeted SFT repair (2,202 steps, about 3.3 GPU-hours, 5x higher learning rate) lifts the 1B's valid emission on 269 corpus rows from 0.100 to 0.959, vs. 0.926 for the 600M.
- On 238 unseen prompts the repaired 1B passes 0.536 vs. 0.428 for the 600M (p = 0.004), but both models over-trigger on negative prompts.
Why it matters
Tool use is a headline capability claim for small language models, and lenient keyword-matching metrics can make a model that never emits a working tool call look as good as one that does. Here the lenient score was nearly a tie (0.660 vs. 0.650), while the strict check showed 6/6 against 0/6. The paper also shows the failure can be diagnosed to a specific cause, a collapsed probability on the <|tool_call|> token, and fixed cheaply: about 3.3 GPU-hours with three orders of magnitude fewer tokens than the failed phase.
Who it affects
Teams that train or evaluate small language models for tool use, and anyone who reads tool-use benchmark numbers for such models. The case study is a pair of Spanish security language models, so the direct audience is people building small, domain-specific models with special tool-call tokens.
How to use it
The authors propose a ladder of strict diagnostics that costs minutes of CPU time: verbatim-reproduction checks on training examples, a first-token probe on the trigger token, and embedding-drift checks. If a model fails, the paper's repair recipe is a diverse corpus, a 5x higher learning rate and 2,202 steps, about 3.3 GPU-hours in their case. The authors say the ladder should gate tool-use claims on small models.
How solid is it
The evidence comes from one matched pair of models, and the source says nothing about results generalizing beyond these two models. The sample sizes are small for the verbatim checks (6 examples). The repaired 1B's edge on 238 unseen prompts (0.536 vs. 0.428) is reported with p = 0.004. What is available is the abstract on a Hugging Face papers page. Numbers here are taken as the authors report them.
Risks and caveats
Both models over-trigger: they rarely answer negative prompts without a call (0.09 for the 600M, 0.17 for the repaired 1B). The authors call the benefit of a diverse corpus for suppression a hypothesis because of seed sensitivity, though all repair configurations installed the format. The source does not say the models were released. No authors or institutions are named in the source.
“Keyword-matching benchmarks can credit small models for tool use they never perform.”
— Opening sentence of the paper's abstract