Microsoft's ThinkingBox: Claude Opus 5.5 passes all 20 runs on just 241 of 507 stateful agent tasks
Microsoft and Hugging Face released ThinkingBox, a benchmark that grades AI agents on the database state they leave behind rather than their responses. Across 507 workflows run 20 times each, Claude Opus 5.5 leads at 67.16% pass@1 but passes all 20 attempts on only 241 tasks.
Microsoft and Hugging Face have released ThinkingBox, a benchmark that grades AI agents on the backend state and side effects they leave behind, not on their final responses or tool calls. It covers 507 stateful business workflows, each run 20 times per model. It is available through Hugging Face and can be run via OpenEnv.
The benchmark is described in a joint blog post and an accompanying paper. All results below are the authors' own and have not been independently replicated.
How it works
Each task runs against isolated MCP tool sessions. Every attempt starts from an identical clean backend, and the grader checks the terminal database state with executable checks. The 507 tasks span five domains: retail (98), auto insurance (100), travel (104), neobank (104) and consulting (101).
Three metrics are reported:
- pass@1: share of all attempts that succeeded.
- pass@20: share of tasks solved at least once in 20 tries.
- Observed 20/20: the literal count of tasks passed on all 20 attempts, with no estimator or smoothing.
Tool calls hide failures
In a common-set ablation of 121,680 valid trials across 12 models, 79,853 attempts failed the executable checks. Of those failures, 67.24% still terminated cleanly, invoked a state-changing tool and reported no final tool error.
The checks found wrong field values in 77.61% of failures, unintended extra effects in 43.30%, and missing required effects in 25.36%. These categories overlap.
The post's example is a retail agent that made nine well-formed tool calls and correctly read the refund policy. It then set a ticket to "solved" when the required end state was "hold" because a carrier exception remained open.
Results
Overall task-weighted pass@1, selected models:
| Model | Overall pass@1 (%) |
|---|---|
| Claude Opus 5.5 | 67.16 |
| Claude Opus 5 | 66.50 |
| GPT-5.4 | 65.36 |
| GPT-5.6 Sol | 61.91 |
| Claude Sonnet 4.6 | 59.19 |
| GPT-6 Astra | 58.31 |
| Kimi-K3 (open-weight) | 57.37 |
| Qwen3.8-27B (open-weight) | 51.70 |
| DeepSeek-V4-Pro (open-weight) | 43.26 |
| Grok-4.3 | 14.38 |
| Mistral-Large-3 | 4.66 |
Domain variance is large. Claude Opus 4.6 scores 68.62% on retail but 8.30% on auto insurance.
Consistency
GPT-6 Astra retains 78% of its pass@1 rate across 20 repeats. Claude Opus 5.5 and Claude Opus 5 each retain 71%. GLM-5.1, Kimi-K2.6 and DeepSeek-V4-Pro each keep about 8%.
Kimi-K3 has the broadest coverage. It solves 476 of 507 tasks (93.89%) at least once, with only 31 tasks defeating it entirely. It passes all 20 attempts on just 68 tasks (13.41%).
Claude Opus 5 solves 79.09% of tasks at least once, with 106 defeating it entirely. It passes all 20 attempts on 241 tasks (47.53%).
Claude Opus 5.5 beats Opus 5 on pass@1 and solves more tasks at least once. It passes the same 241 tasks on all 20 attempts.
What this means
Most agent evaluations score the transcript or the tool-call sequence. ThinkingBox's data suggests that approach misses the majority of real failures: two-thirds of failed attempts looked clean from the outside.
The Opus 5 to 5.5 comparison is the sharpest finding. A 0.66-point pass@1 gain produced no additional fully reliable tasks. Teams should treat headline accuracy and repeatability as separate selection criteria.
The Kimi-K3 versus Opus 5 split, 75 more tasks solved at least once against 173 more solved every time, matters for deployment. Pass@20 suggests what a model might do with retries and verification. Observed 20/20 suggests what it can do unsupervised on records that matter.
Caveats: the benchmark and results come from Microsoft, and the task suite is synthetic business workflows. Pricing and context-window figures are not part of this release.
Related Articles
Mercor study: AI models beat 12 licensed CPAs on simplified accounting tasks, but top score on full APEX benchmark is 61
A Mercor study found AI models beat 12 licensed CPAs on simplified tasks from the APEX Accounting Benchmark. On the full 160-task benchmark, the top model, Claude Opus 5.5, meets 61.8% of grading criteria, and no model fully solved almost 60% of tasks.
Microsoft's MAI-Transcribe-2-Streaming returns first results in ~100 ms across 60 languages
Microsoft AI released MAI-Transcribe-2-Streaming, a real-time transcription model covering 60 languages with first partial results in just over 100 milliseconds. It also launched two text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, aimed at voice agents.
Microsoft Restructures Copilot Into Three Apps, Adds Autopilot Agent and Usage-Based Billing
Microsoft is overhauling Copilot with three distinct sections—Home, Code, and Autopilot—headlined by a proactive business agent built on OpenClaw. The company is also replacing flat-rate pricing with usage-based billing for its agent and automation tools.
Microsoft Merges Coding and Productivity Copilot Into Single App to Counter Anthropic
Microsoft launched an updated Copilot app that merges coding, productivity tasks, and custom agent creation into three tabs — Cowork, Code, and Autopilot. The company is shifting to usage-based pricing as it tries to close the gap with Anthropic's Claude in enterprise AI adoption.
Comments
Loading...