TL;DR
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Microsoft and Hugging Face have made ThinkingBox available through Hugging Face. The benchmark evaluates AI agents on whether they leave the required database state and side effects across 507 business workflows, repeated 20 times per task. Its authors report substantial gaps between clean tool execution and passing the checks, though the results come from the benchmark’s own tested setup.
The benchmark authors describe a retail support task in which an agent investigates a delayed $745 appliance order, opens a ticket, and records the timeline. The customer’s account does not qualify for late-delivery compensation under the policy the agent checked. But the delivery carrier exception remains open, while the task requires the ticket to stay on hold pending resolution. The agent marks the ticket solved and replies without answering the customer’s underlying question. The authors say the executable check fails because the ticket status is solved instead of hold.
ThinkingBox runs agents against isolated sessions using MCP tools, then checks the terminal backend state and side effects. In a common-set analysis covering 121,680 valid trials across 12 models, the authors report that 79,853 attempts failed executable checks. Of those failures, 67.24% ended without a final tool error despite invoking a state-changing tool. Among the failures, checks found wrong field values in 77.61%, unintended extra effects in 43.30%, and missing required effects in 25.36%; these categories overlap.
The release describes several ways to summarize performance. Pass@1 measures the share of attempts that succeed; pass@20 asks whether a task succeeded at least once in 20 runs; and observed 20/20 counts tasks that passed every recorded run. The authors report an overall pass@1 of 67.16% for Claude Opus 5.5 and 57.37% for Kimi-K3, the strongest open-weight model in their table. They say Kimi-K3 scored within a point of GPT-6 Astra, but the provided material does not supply the full table’s uncertainty estimates.
AI AGENT BENCHMARK · HUGGING FACE
The Agent Said It Was Done. The Database Disagreed.
ThinkingBox checks what an agent leaves behind: the records and side effects a business workflow requires. Across 507 workflows and 20 runs per task, a convincing reply is only the beginning of the test.
“A tool call is not an outcome.”
ThinkingBox release · Microsoft and Hugging FaceA delayed $745 appliance order. A carrier exception still open. The required ticket status: hold. The agent recorded solved.
THE GAP
Fluent interaction. Failed outcome.
In the authors’ retail example, the agent investigates the case, opens a ticket, and records a timeline. The backend still says the task was not completed correctly.
Handled the case
The customer asked about a delayed $745 appliance order. The agent checked the compensation policy, found the customer did not qualify, opened a ticket, and recorded the timeline.
Closed it too soon
The carrier exception remained open, so the ticket needed to stay on hold. The agent marked it solved and replied without answering the underlying question.
WHAT FAILURES LOOKED LIKE
Valid tools can still leave broken state.
Of 121,680 valid trials across 12 models, the authors report 79,853 attempts failed executable checks.
Wrong field values
A record had a value different from what the task required.
Unintended extra effects
The run changed something the task did not call for.
Missing required effects
A necessary update or side effect never happened.
Failure categories overlap. Among failed attempts, 67.24% ended without a final tool error despite invoking a state-changing tool.
READING THE SCORES
One success is not consistency.
ThinkingBox reports multiple views of success. Each answers a different question about performance across repeated runs.
Three ways to count
Pass rates describe outcomes observed within this benchmark setup.
The bars show run coverage, not model scores.
Reported overall pass@1
Authors’ results from the supplied table excerpt.
The authors say Kimi-K3 scored within one point of GPT-6 Astra. The supplied material does not include uncertainty estimates for the full table.
HOW THINKINGBOX CHECKS
From isolated run to verified state.
The benchmark focuses on the system’s terminal backend state and side effects after an agent acts.
Start clean
Each workflow begins in an isolated backend session.
Agent acts
The agent uses MCP tools to carry out the task.
Run repeats
Each task runs 20 times to observe consistency.
Check evidence
Executable checks inspect final records and side effects.
TRACE THE EVIDENCE
A trajectory is a claim. State is the evidence.
A useful evaluation follows the workflow all the way to the recorded outcome.
LIMITS & PRACTICAL USE
Benchmark evidence has a boundary.
The results are the authors’ findings on their tested workflows and setup.
What it can show
Where agents miss required state changes, create extra effects, or perform inconsistently across repeated workflow trials.
What it cannot settle
Whether these scores predict performance in every live business system, policy setting, or changing operational environment.
What the supplied excerpt omits
Full task specifications, model configurations, score uncertainty estimates, and a publication date for the release.
A practical next step
Run relevant workflows against your own requirements and inspect both passing and failing runs. OpenEnv is named as a way to run the benchmark.
KEY QUESTIONS
What to take away
What does ThinkingBox measure?
Whether an agent leaves backend records and side effects in the state required by a workflow, including how often it succeeds over repeated runs.
How broad is the benchmark?
507 stateful workflows, each run 20 times from a clean backend, across five named business domains.
Which models led the reported results?
Claude Opus 5.5 scored 67.16% overall pass@1; Kimi-K3 led the listed open-weight models at 57.37%.
Can a clean tool trace still fail?
Yes. In the retail example, the ticket was marked solved when the required status was hold, so the backend check failed.
Why Backend Checks Change Agent Scores
For businesses using agents to handle refunds, support tickets, insurance claims, or bookings, a fluent answer does not establish that the requested work was completed. A ticket may be closed too early, a record may contain the wrong value, or an unintended change may have been made. Backend checks expose those outcomes even when the interaction appears orderly.
Repeated trials address a separate deployment question: whether an agent can perform reliably, not merely whether it can succeed once. A high pass@1 score still leaves room for failures, while a task that passes all 20 observed runs offers stronger evidence within this test setup. The benchmark can help teams compare models and identify workflow weaknesses, but its scores alone do not establish performance in a company’s own systems.
As an affiliate, we earn on qualifying purchases.
From Tool Calls to Recorded Outcomes
The ThinkingBox authors frame their work around a distinction between an agent’s actions and the resulting system state. Tool-call validity and final-response quality are indirect signals: an agent may use tools correctly while leaving a record that fails the task’s requirements. ThinkingBox instead evaluates the records and side effects after each run.
The benchmark covers five areas named in the release: retail, auto insurance, travel, neobanking, and consulting. Each task is run repeatedly from a clean backend. The source says the release is based on the authors’ paper and that the benchmark can be run through OpenEnv; it does not provide a publication date in the supplied material.
“A tool call is not an outcome.”
— Microsoft and Hugging Face, in the ThinkingBox release
business process automation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Limits of the Reported Results
The supplied release excerpt does not specify when ThinkingBox became available, nor does it include all details of the paper’s evaluation setup, such as the full task specifications, model configurations, or uncertainty estimates for the reported scores. The reported failures and rankings are the authors’ results on this benchmark; they do not show how the same models would perform across every live business system or under different operating conditions.
The source also does not establish whether passing 20 trials predicts long-term reliability. Twenty clean runs provide a bounded observation, and real deployments may include changing records, unusual customer requests, integrations, or policies not represented in the tested workflows.
As an affiliate, we earn on qualifying purchases.
Running the Benchmark Independently
The release says readers can run ThinkingBox through OpenEnv using isolated MCP tool sessions. That gives researchers and developers a way to inspect the benchmark tasks and compare agent outcomes against executable state checks. The supplied material does not name a future release date, independent replication, or a planned next milestone.
For teams considering deployment, the practical next step is to test relevant workflows against their own requirements and inspect both successful and failed runs. The benchmark’s authors present repeated state checks as a way to measure consistency; whether that predicts performance in a particular organization remains to be established.
As an affiliate, we earn on qualifying purchases.
Where I land
I see ThinkingBox as a useful response to a real evaluation gap: a polished answer and a sequence of valid tool calls do not prove that an agent completed a workflow correctly. The delayed-order example makes that gap concrete, and the repeated trials add evidence about consistency that a single successful run cannot provide.
The strongest counterargument is that benchmark scores can look more decisive than they are. A fixed set of 507 workflows and 20 runs per task cannot capture every live system, policy exception, or changing record. I would put more weight on the results if independent teams reproduced them and if the benchmark’s task coverage and scoring held up against real deployment outcomes. Evidence that these scores reliably predicted performance in varied production settings would strengthen my assessment; evidence of large gaps would narrow it.
Key Questions
What does ThinkingBox measure?
It checks whether an AI agent leaves the backend records and side effects in the state required by a business workflow. It also measures how often the agent succeeds across repeated runs.
How many workflows does the benchmark include?
The release reports 507 stateful business workflows, with each task run 20 times from a clean backend. It names retail, auto insurance, travel, neobanking, and consulting as the domains.
Which models scored highest in the reported results?
In the authors’ overall pass@1 table, Claude Opus 5.5 scored 67.16%. Kimi-K3 led the listed open-weight models at 57.37%. These are benchmark results, not guarantees of performance in other settings.
Can a clean tool trace still fail the benchmark?
Yes. In the retail example, the agent used tools and documented the case but marked the ticket solved when the required status was hold. The backend check therefore failed the task.
Source: Hugging Face
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
