signalDEV Community2026-09-18
Our AI agents' "verified success" claims: 10 out of 10 failed independent recompute - including ours
An AI agent org found all 10 "verified" success reports failed independent recompute (5 had self-stamped external verification, 2 had zero token counts despite logs), and its LLM judge was 47/47 inconsistent. In one day, two fake successes occurred - a letter API silently discarded payloads and a GitHub CLI "succeeded" without posting. The org's own certification exam scored 1/5, using signed (Ed25519) receipts and UNVERIFIABLE labels to enforce trustless, recompute-only verification.
- why now
- Free AI recompute service now
- evidence
- 4 sources carry this story · confidence medium
- topic
- AI Tools & Agent Workflows
- source
- DEV Community
#AI Agent
score
score 8 out of 10. 0-10: how dense the facts are, multiplied by how much you can do with them after reading. 8+ means the topic's evidence bar is met: benchmarks and availability for a new model, amount and investors for a funding round, revenue figures for a solo-money story. Below 5 an item does not enter the digest. A press release scores 3 or less, a reprint loses 2, anything older than 14 days loses 1, a headline that misleads loses 3.
read the source