signalDEV Community2026-09-27
AI models catch bad code, then cry wolf on the good code
A new Kaggle benchmark, Blog vs Bytecode, tests AI models on 28 data-science snippets with balanced claims. A harness flaw that counted empty responses as wrong made strong models look broken, for example DeepSeek-R1 jumped from 17% to 100% after fixing capture. Frontier models like Gemini 3.1 Pro and Grok 4.20 with reasoning achieve 100% accuracy, but they over-flag clean code, while smaller models like Gemma 4 31B under-flag, catching only 20% of flaws.
- for who
- Anyone evaluating or building AI code review tools, and designers of model benchmarks
- why now
- New Kaggle benchmark reveals frontier AI models over-flag clean code, a current evaluation issue.
- what changes
- They see that current AI models over-flag correct code, and that benchmark infrastructure must audit response capture to avoid false results
- to do
- Use the public Blog vs Bytecode benchmark to test models, and ensure any evaluation harness flags empty responses instead of scoring them wrong
key points
- Benchmark of 28 snippets, 15/13 balanced, tests for over-flagging and under-flagging
- Harness bug made DeepSeek-R1 show 17%, then 100% after fixing empty responses
- Big models over-flag, small under-flag; Grok 4.20 gains 32 points with reasoning
#model evaluation#benchmarking#code review
score
score 9 out of 10. 0-10: how dense the facts are, multiplied by how much you can do with them after reading. 8+ means the topic's evidence bar is met: benchmarks and availability for a new model, amount and investors for a funding round, revenue figures for a solo-money story. Below 5 an item does not enter the digest. A press release scores 3 or less, a reprint loses 2, anything older than 14 days loses 1, a headline that misleads loses 3.
read the source