AI NewsModels & agentsAnnouncement
Microsoft's ThinkingBox benchmark grades AI agents on what they left in the database, over 20 runs per task
Microsoft Research and Hugging Face released ThinkingBox, a benchmark that runs 507 stateful business workflows 20 times per model and grades each attempt on the records the agent actually left in the database.

Image: Microsoft and Hugging Face
Why it mattersAn agent that passes a check once can fail it on the next run, so a team deploying one needs to grade on database state, repeat each case many times, and design the workflow around the 20 of 20 rate instead of the single-attempt score.
Any team that has shipped an agent to touch real records knows the gap: the agent returns a tidy "done", and the database says otherwise.
Microsoft Research and Hugging Face published a benchmark today that scores agents only on what they left behind, and runs every task 20 times to see how often the result holds. It is called ThinkingBox. It covers 507 stateful workflows across retail, auto insurance, travel, neobank, and consulting, each with a clean database, a simulated user that holds back private details until asked, and executable checks that compare the final database state against the required end state. Tuhin Kundu at Microsoft wrote it up on the Hugging Face blog on 3 October 2026.
Where the state check caught the agents
In a common-set ablation across 121,680 valid trials and 12 LLM models, Microsoft reports that 79,853 attempts failed the executable checks. Of those failures, 67.24% still terminated cleanly, invoked a state-changing tool, and reported no final tool error. The side-effect extractor found wrong field values in 77.61% of them, extra effects in 43.30%, and missing required effects in 25.36%.
Repetition changes the ranking
Every task runs 20 independent times from an identical backend. The benchmark reports pass@1, pass@20, and observed 20/20 (tasks passing all 20 recorded attempts). Kimi-K3 solves 93.89% of the tasks at least once, the broadest coverage in the test, and only 13.41% on all 20 attempts. Claude Opus 5 solves fewer tasks at least once (79.09%) but reaches 20 of 20 on 47.53%. For work that touches real records, pass@1 and pass@20 are the wrong columns to read.
Dependable is a different price
Microsoft prices the campaign at undiscounted OpenRouter list rates. By cost per successful task attempt, the Pareto frontier is GPT-5.6 Sol at $0.127, GPT-5.4 at $0.131, and Claude Opus 5.5 at $0.276. By cost per task passed on all 20 runs, the ranking changes: GPT-5.4 is cheapest at $6.80 (128 dependable tasks), GPT-6 Astra reaches $7.45 (231), and Claude Opus 5.5 sits at $7.80 (241). Microsoft calls the figures a comparative efficiency index.
Why the failures happen
Each failed trajectory gets one deterministic signature. The shares: 79.9% tool usage, 10.3% wrong state updates, 7.0% incomplete user resolutions, 2.9% no state-changing action. The post calls this a retry and error-recovery problem before it is a model problem.
What a team can run
ThinkingBox code is MIT, the benchmark data is CDLA-Permissive-2.0, and the OpenEnv adapter is BSD-3-Clause. Of the 507 tasks, 477 are graded on state alone and 30 add a narrow binary response rubric. Every task runs in an isolated MCP session with freshly initialised state, so two attempts of the same task never share a row or a cached tool result.
Reveneau grades its eval suite with Jev, TypeSafe AI's decision model, measured against our previous grader. A benchmark that scores on terminal state and repeats every case 20 times is the same discipline, applied to an agent instead of a code change. Microsoft is direct about what to do: treat the 20 of 20 rate as a design input, check terminal state before you commit, classify tool errors so retries target the recoverable ones, cut the tool surface to what the workflow needs, and require human approval on changes the system cannot cheaply reverse.
Source
- Hugging Face blog post, by Tuhin Kundu (Microsoft): The Agent Said It Was Done. The Database Disagreed.
- Benchmark dataset: microsoft/ThinkingBox-Bench
- Paper: One Success Isn't Reliability: ThinkingBox, a Sandbox and Benchmark for Agents in Stateful Business Workflows
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.


