signalAI热榜2026-10-04
Microsoft and Hugging Face Release ThinkingBox for Agent Evaluation with Database End States and 20 Repetitions
Microsoft and Hugging Face released ThinkingBox, an agent sandbox and benchmark that evaluates agents across 507 stateful business workflows, each run 20 times, scoring on terminal database states and side effects. Available via OpenEnv on Hugging Face, the benchmark shows that most models fail to retain reliability across repetitions, with only three models retaining over 70% of their single-trial scores.
- for who
- AI engineers and developers evaluating agent reliability in stateful business workflows.
- why now
- Microsoft's ThinkingBox benchmark, now on Hugging Face, evaluates agents via database states and 20 repeats.
- what changes
- They can now assess agents based on database end states and consistency across repeated runs, revealing the gap between single-trial success and dependable performance.
- to do
- Run ThinkingBox-Bench via OpenEnv on Hugging Face to evaluate your own agents on 507 workflows.
key points
- 507 workflows, each run 20 times, scored on final database states
- Only GPT-6 Astra, Claude Opus 5.5, and Claude Opus 5 retained over 70% of pass@1
- Kimi-K3 has broadest coverage but only 13.41% of tasks consistently solved
#agent evaluation#benchmark#Microsoft#Hugging Face#workflow reliability
score
score 8 out of 10. 0-10: how dense the facts are, multiplied by how much you can do with them after reading. 8+ means the topic's evidence bar is met: benchmarks and availability for a new model, amount and investors for a funding round, revenue figures for a solo-money story. Below 5 an item does not enter the digest. A press release scores 3 or less, a reprint loses 2, anything older than 14 days loses 1, a headline that misleads loses 3.
read the source