ActiveSaddler adapts training scenarios to improve agent harnesses

A paper on Hugging Face introduces ActiveSaddler, a method for automated harness optimization. In this setting, an LLM agent's harness (its prompts, tool interfaces and control logic) is updated iteratively from execution feedback, which can substantially improve the agent.
The authors argue that existing methods mainly optimize how the harness is updated, while largely fixing which training scenarios generate the feedback that drives those updates. Their point is that as the harness evolves, the scenarios most useful for further optimization can change, so the training curriculum should adapt alongside the harness. They frame this missing dimension as an automated curriculum learning problem.
ActiveSaddler models the evolving curriculum as a non-stationary bandit with dynamically instantiated optimization targets. It abstracts recurring failures into reusable failure-pattern arms. It then estimates the potential learning progress from further targeting each pattern, and adaptively balances revisiting known weaknesses with exploring unseen scenarios for new ones. Optimization outcomes keep updating both the set of discovered failure patterns and their priorities, so the curriculum co-evolves with the harness.
On GAIA2 and Terminal-Bench 2.0, the authors report that ActiveSaddler consistently discovers stronger harnesses. Test Pass@1 improves by 4.4 percentage points on GAIA2 and 7.5 percentage points on Terminal-Bench 2.0, compared with the same harness optimizer using a scenario order fixed before optimization.
Ablations show that the gains depend on three things: dynamically constructing optimization targets, estimating their evolving utility, and balancing continued optimization with new failure discovery. The authors conclude that automated curriculum learning is a new crucial optimization dimension for harness optimization.
Key facts
- ActiveSaddler treats the choice of training scenarios in agent-harness optimization as an automated curriculum learning problem, modelled as a non-stationary bandit.
- It groups recurring failures into reusable failure-pattern arms and estimates the learning progress to be gained from targeting each one.
- It balances revisiting known weaknesses with exploring unseen scenarios to find new failures.
- Test Pass@1 improves by 4.4 percentage points on GAIA2 and 7.5 percentage points on Terminal-Bench 2.0, versus the same optimizer with a scenario order fixed before optimization.
- Ablations indicate the gains depend on dynamic target construction, utility estimation and the balance between continued optimization and new failure discovery.
Why it matters
Automated harness optimization rewrites an agent's prompts, tool interfaces and control logic from execution feedback. The authors say existing methods focus on how the harness is updated and largely fix which scenarios supply the feedback. ActiveSaddler adds a second axis: as the harness changes, the most useful training scenarios change too, so the curriculum should adapt. The authors present this as a new crucial optimization dimension.
Who it affects
The work is aimed at people building and tuning LLM agents with automated harness optimization, where prompts, tool interfaces and control logic are updated from execution feedback. Teams evaluating agents on GAIA2 or Terminal-Bench 2.0 are the ones the reported results speak to most directly.
How to use it
The abstract describes the approach rather than a recipe. In outline: abstract recurring failures into failure-pattern arms, estimate the potential learning progress from targeting each pattern, and balance revisiting known weaknesses with exploring unseen scenarios, updating the set of patterns and their priorities after each optimization outcome. No code release, dataset release, or publication venue is mentioned in the source.
How solid is it
The evidence is the authors' own experiments on two benchmarks, GAIA2 and Terminal-Bench 2.0, with gains of 4.4 and 7.5 percentage points in test Pass@1 over the same optimizer using a fixed scenario order. Ablations are reported to support the design choices. No absolute Pass@1 scores or baseline values are given, only the percentage-point gains. The text available is the paper's abstract.
Risks and caveats
The abstract does not say which LLM or which harness optimizer was used, and it does not name the fixed-order baseline optimizer. No compute cost, number of optimization iterations, or number of failure patterns is given. The gains are measured against a fixed scenario order, and the claim that curriculum learning is a crucial dimension is the authors' own conclusion from two benchmarks.
“As the harness evolves, the scenarios most useful for further optimization can change, suggesting that the training curriculum itself should adapt alongside the harness.”
— ActiveSaddler paper abstract