Agents

Microsoft releases ThinkingBox to test AI agent databases

Microsoft has launched ThinkingBox on Hugging Face, a new benchmark that evaluates AI agents based on actual database changes rather than text outputs to measure real-world reliability.

Hugging Face Blog3 days agoAgents
Image: Hugging Face Blog

Microsoft, in partnership with Hugging Face, has released ThinkingBox, an evaluation framework that grades AI agents on the actual database records and side effects they leave behind. Traditional benchmarks rely on text responses or tool calls as proxies. However, Microsoft's research shows that agents frequently sound correct while failing to update backend systems. In an ablation of 121,680 trials across 12 models, 79,853 attempts failed database checks. Strikingly, 67.24% of those failures terminated cleanly with no tool errors, yet checks revealed incorrect field values in 77.61%, unintended extra effects in 43.30%, and missing effects in 25.36%.

ThinkingBox-Bench tests agents across 507 stateful workflows, running each task 20 times to measure consistency. The results expose a massive gap between single-attempt success (pass@1) and repeated reliability. Kimi-K3 solved 93.89% of tasks at least once and led retail with an 82.24% pass@1, but completed all 20 runs on just 13.41% of tasks. Conversely, Claude Opus 5 completed 47.53% of tasks consistently across all 20 attempts. Claude Opus 5.5 led overall with a 67.16% pass@1, retaining 71% of its score over 20 repeats, while GPT-6-Astra retained 78%. Models like GLM-5.1, Kimi-K2.6, and DeepSeek-V4-Pro retained only about 8% of their scores.

The benchmark also calculates costs based on September 20, 2026 pricing. GPT-5.6 Sol had the lowest cost per single success at $0.127, while GPT-5.4 cost $0.131 with a 65.36% pass@1. For dependable tasks (passing 20 out of 20 runs), GPT-5.4 was cheapest at $6.80, followed by GPT-6-Astra at $7.45, and Claude Opus 5.5 at $7.80. Claude Opus 5 cost $13.30 for the same consistency. For practitioners, the data shows that roughly four in five failures stem from tool handling rather than reasoning, meaning developers should verify terminal database states before committing transactions.

This is our own summary of reporting by Hugging Face Blog

More in Agents