25 signals
HOlO V1 IS LIVEone ranked AI digest a day, scored in publicREAD HOW IT WORKS →HOlO V2 STARTSyour X account, your signals, every day
signalHugging Face Blog2026-10-04

The Agent Said It Was Done. The Database Disagreed.

ThinkingBox, a joint benchmark from Microsoft and Hugging Face, evaluates AI agents by grading the terminal backend state and side effects they leave behind, across 507 stateful workflows run 20 times each. In 121,680 trials across 12 models, 79,853 failed executable checks, with 67.24% of failures terminating cleanly while still leaving wrong state. The benchmark reveals that tool-call validity is not a proxy for correct outcomes.

for who
Developers and researchers building or evaluating AI agents that interact with databases and stateful systems.
why now
New benchmark reveals agents often fail real database checks, vital for evaluating AI reliability now.
what changes
Agent evaluation shifts from trusting tool-call traces to verifying the actual end state, exposing silent failures and improving reliability.
to do
Run the ThinkingBox benchmark yourself via OpenEnv and adopt its state-based grading to catch failures that clean tool calls miss.
key points
  • 507 workflows run 20 times each, grading terminal backend state
  • 79,853 of 121,680 trials failed despite clean tool calls
  • Microsoft and Hugging Face release ThinkingBox via OpenEnv
#agent evaluation#thinkingbox#database state check#microsoft#hugging face
score
score 8 out of 10. 0-10: how dense the facts are, multiplied by how much you can do with them after reading. 8+ means the topic's evidence bar is met: benchmarks and availability for a new model, amount and investors for a funding round, revenue figures for a solo-money story. Below 5 an item does not enter the digest. A press release scores 3 or less, a reprint loses 2, anything older than 14 days loses 1, a headline that misleads loses 3.
read the source