Microsoft and Hugging Face present ThinkingBox, a benchmark that grades agents on database state

Microsoft and Hugging Face present ThinkingBox, a benchmark that grades agents on database state

Microsoft and Hugging Face have published a joint blog post presenting ThinkingBox, an agent sandbox, and ThinkingBox-Bench, a benchmark of 507 stateful business workflows. Each workflow is run 20 times against each LLM, and the agent is graded on the terminal backend state and side effects it leaves behind, not on its tool calls or its final reply.

The post opens with an example adapted from a benchmark task. A customer's $745 kitchen appliance is stuck in a courier exception at a Nashville distribution center, fifteen days past its estimated delivery date. The agent makes nine well-formed tool calls: it pulls the order, checks tracking, looks up the customer profile, searches the refund policy twice, confirms no ticket exists, opens one, documents the timeline, and reads the policy correctly (her account segment does not qualify for late-delivery compensation). Then it closes the ticket as resolved and replies, "Since your query is resolved, is there anything I may assist you with?" The carrier exception is still open, so the required end state was "on hold, pending resolution", and the customer never got a real answer to her question. A grader looking at tool calls, or at whether the agent wrote to the database, would see nothing wrong. The one failing executable check is the ticket status: it is "solved" where the required end state is "hold".

The authors argue that final responses and valid tool calls are only proxies. In a common-set ablation covering 121,680 valid trials across 12 LLMs, 79,853 attempts failed the executable checks. Of those failures, 67.24% still terminated cleanly, invoked a state-changing tool and reported no final tool error. Among them, executable checks found wrong field values in 77.61%, unintended extra effects in 43.30% and missing required effects in 25.36%; these findings overlap.

Because one success is not reliability, every task runs 20 independent times from an identical clean backend. The post reports three kinds of number; the one it leans on is "observed 20/20", the literal count of tasks, out of 507, passed on all 20 attempts, with no estimator and no smoothing. On the familiar single-attempt pass@1 view, Claude Opus 5.5 leads overall at 67.16%, two-thirds of a point above Claude Opus 5 (66.50%). Kimi-K3 is the strongest open-weights model, within a point of GPT-6-Astra. Domain matters: Claude Opus 4.6 scores 68.62% on retail but 8.30% on auto insurance, and across the models in the paper's Table 2 retail averages 59.52% pass@1 against 33.83% for auto insurance.

Repetition changes the picture. GPT-6 Astra keeps 78% of its single-attempt rate over 20 repeats, and Claude Opus 5.5 and Claude Opus 5 each keep 71%. GLM-5.1, Kimi-K2.6 and DeepSeek-V4-Pro each keep about 8%. Kimi-K3 has the broadest coverage: it solves 476 of 507 tasks (93.89%) at least once, with only 31 defeating it entirely, the lowest count in the field, and it leads retail at 82.24% pass@1. But only 68 tasks (13.41%) succeed in all 20 attempts. Claude Opus 5 is the reverse: 79.09% solved at least once (106 tasks defeat it entirely), yet 47.53% of the benchmark completed on every attempt. Opus 5.5 solves more tasks at least once than Opus 5 but passes exactly the same 241 tasks on all 20 attempts. Kimi-K3 solves 75 more tasks at least once than Opus 5; Opus 5 solves 173 more tasks consistently. The authors' advice: for work that touches real records, pass@20 is the wrong column to look at.

The cost analysis prices each model's recorded token usage from its full 507 x 20 campaign at undiscounted list rates on OpenRouter+, reversing promotional discounts and excluding endpoints that declare quantization. Cost per successful task attempt is the cost of 507 attempts divided by (507 x pass@1); for GPT-5.4 that is $43.49 at 65.36% pass@1, or $0.131. Three models sit on the cost frontier. GPT-5.6 Sol is cheapest per success at $0.127; GPT-5.4 adds 3.45 percentage points of pass@1 for $0.004 more per success; Claude Opus 5.5 adds another 1.80 points at $0.276 per success. Claude Opus 5, at $0.475 per success and 66.50% pass@1, is both costlier and less accurate than Opus 5.5.

A second measure, cost per dependable task, divides the cost of the full 20-run campaign by the tasks passed on all 20 attempts. GPT-5.4 is cheapest at $6.80 but only 128 tasks meet the bar. GPT-6 Astra reaches 231 tasks at $7.45 (a $1,720.60 campaign, 20 x $86.03), and Claude Opus 5.5 reaches the joint-highest 241 at $7.80. Claude Opus 5 also passes 241 but at $13.30, so Opus 5.5 dominates it. GPT-5.6 Sol, cheapest per single success, costs $9.76 per dependable task. In the authors' words, the cheapest way to get a right answer is not the cheapest way to get a dependable one.

On failure signatures, each failed trace gets one deterministic diagnostic label, and the authors report that roughly four in five failures are tool handling, not reasoning. Agents usually get far enough to attempt the workflow, then fail to recover from tool errors, failed preconditions or empty lookups, which the authors call a retry and error-recovery problem before a model problem. They suggest checking terminal state before committing, classifying tool and system errors so retries target recoverable ones, cutting the tool surface to what the workflow needs, and requiring human approval for changes that cannot cheaply be reversed. They say they have not measured the lift from any of these on this benchmark.

As for design, each task defines a starting backend state, a user goal, the available MCP tools, the domain policy and executable checks over the terminal state. A simulated user holds private context, such as a booking reference, a preference or a date of birth, and releases it only when asked. Every attempt gets an isolated MCP session with freshly initialized state, so two attempts never share a database row or cached tool state. At the end a side-effect extractor derives what changed, and deterministic judges compare it with the required end state, accepting any trajectory that reaches the right outcome and rejecting wrong, missing or extra effects. A narrow binary rubric question covers requirements with no clean database value. The post says the benchmark can be run through OpenEnv.

Key facts

  • ThinkingBox-Bench has 507 stateful business workflows, each run 20 times per model, graded on the terminal database state and side effects rather than tool calls or replies.
  • In a 12-model ablation of 121,680 valid trials, 79,853 attempts failed the checks; 67.24% of those failures still ended cleanly with no final tool error.
  • Claude Opus 5.5 leads pass@1 at 67.16% (Claude Opus 5: 66.50%), but both pass the same 241 tasks on all 20 attempts; GPT-6 Astra keeps 78% of its single-attempt rate over repeats, Opus 5.5 and Opus 5 keep 71%, and GLM-5.1, Kimi-K2.6 and DeepSeek-V4-Pro keep about 8%.
  • Kimi-K3 solves 476 of 507 tasks at least once but only 68 (13.41%) on every attempt; GPT-5.6 Sol is cheapest per success ($0.127) yet costs $9.76 per dependable task, against $6.80 for GPT-5.4 (128 tasks) and $7.80 for Claude Opus 5.5 (241 tasks).
  • The authors say roughly four in five failures are tool handling, not reasoning, and that suggested fixes such as error classification and human approval have not been measured on this benchmark.

Why it matters

Most agent evaluation looks at what the agent says or which tools it calls. The ThinkingBox authors argue that both are only proxies, and that the records left in the database are the evidence. Their opening example makes the point: nine well-formed tool calls, a correct reading of the policy, and still a ticket closed as solved when it should have been on hold. The numbers back the framing. Of 79,853 failed attempts in the ablation, 67.24% ended cleanly with a state-changing tool call and no final tool error, so an agent that looks fine can still leave the wrong data behind. Running every task 20 times also shows how far a single-attempt score can mislead: the share of pass@1 retained over repeats ranges from 78% for GPT-6 Astra to about 8% for three other models.

Who it affects

Teams choosing a model to run business workflows that change real records, such as retail support or insurance, are the direct audience. The authors say that for such work pass@20 is the wrong column to look at, and that a newer model with a higher headline score does not necessarily bring more dependability: Opus 5.5 and Opus 5 pass the same 241 tasks on all 20 attempts. The results also vary by domain. Claude Opus 4.6 scores 68.62% on retail and 8.30% on auto insurance, so a ranking in one domain may not carry over to another. Model builders and evaluators are affected too, since the benchmark grades outcomes rather than transcripts.

How to use it

The post says the benchmark can be run through OpenEnv; the instructions themselves are not in the text available here. The practical advice it gives is to treat the 20/20 rate as a design input, not a verdict. Check the terminal state before committing, not the model's summary of it. Classify tool and system errors so retries target the recoverable ones. Cut the tool surface to what the workflow needs. Require human approval on changes you cannot cheaply reverse. On cost, the authors offer two measures. Cost per successful task attempt is a comparative efficiency index, not an invoice and not the price of serving one production request. Cost per dependable task divides the 20-run campaign cost by the tasks passed on all 20 attempts, which prices consistency. Both are estimates at undiscounted OpenRouter+ list rates.

How solid is it

This is a joint Microsoft and Hugging Face blog post that summarises their ThinkingBox paper, so the results are the authors' own. The method is described in concrete terms: isolated MCP sessions with freshly initialized state, deterministic judges over extracted side effects, 20 trials per task, and a plain count of tasks passing 20 out of 20 with no estimator or smoothing. Standard errors for the pass@1 estimates are said to be in Table 4 of the paper, but the tables and figures are not reproduced in the text used here, so some details, such as the full per-model table and the exact failure-signature shares, could not be checked. The failure-signature shares are described as unweighted averages of per-model shares and observable labels, not unique causal explanations. Cost figures are estimates, not actual cloud bills.

Risks and caveats

The authors say the cost per successful attempt prices single successes, not consistency, and that the cost figures are estimates from list prices, not actual bills. The failure-signature shares are observable labels and not causal explanations, and the claim that roughly four in five failures are tool handling comes from an ablation reported in Table 5 of the paper, averaged across models. The suggested fixes (checking terminal state, error classification, a smaller tool surface, human approval) are recommendations; the authors say they have not measured their lift on this benchmark. The opening customer case is adapted from a benchmark task and is an illustration. Figure 3 shows twelve of the eighteen models tested; six below 33% pass@1 are omitted. The 20/20 count is a strict bar, so it favors models that are steady on a narrower set of tasks.

“A trajectory is a claim. Database state is the evidence. Repetition is the trust test.”

— ThinkingBox authors, Microsoft and Hugging Face blog post