A customer is waiting for a $745 kitchen appliance stuck in a Nashville distribution center. An AI agent performs nine tool calls, checks the order, reads the refund policy, and closes the ticket as "solved." However, the database disagrees: the carrier exception remains open, and the customer never received a real answer.

The Gap Between Claims and Evidence

This case study is at the heart of ThinkingBox, a new benchmark developed by the Microsoft Copilot Studio team in partnership with Toloka and academic collaborators. ThinkingBox doesn't judge agents by their responses, but by the terminal backend state and side effects they leave behind.

In a study of 121,680 trials across 12 LLM models, 67.24% of failures terminated cleanly, invoking state-changing tools without reporting errors, despite the outcome being wrong. Executable checks found incorrect field values in 77.61% of these failures and unintended side effects in 43.30%.

Consistency: The True Challenge

ThinkingBox introduces the pass@20 metric: how many times an agent correctly completes the same task across 20 independent trials. The results show that capability does not guarantee dependability:

  • Claude Opus 5.5 and GPT-6 Astra retain most of their single-attempt scores when repeated.
  • Kimi-K3, while solving the most tasks at least once (93.89%), achieves a perfect 20/20 score in only 13.41% of cases.

Failure Analysis and Economics

The research finds that roughly 80% of failures stem from tool handling rather than reasoning flaws. Agents often fail to recover from failed preconditions or empty lookups. Economically, GPT-5.4 emerged as the cheapest for "dependable tasks" ($6.80 per 20/20 success), followed by Claude Opus 5.5 at $7.80.

ThinkingBox is now open-sourced on Hugging Face, providing a sandbox environment where developers can verify terminal states before committing changes, rather than relying on the model's summary of its own work.