signalHacker News Show HN2026-09-28
Show HN: Has Anthropic been nerfing their models without disclosure
A new deterministic benchmark, livenerf, tracks whether Anthropic's Claude Opus 5.5 degrades after launch. It runs daily for 30 days, using a panel of 78 calibrated questions, and can detect an accuracy change of about 7.5 points per 10-day window. Early data shows lower effort reduces output tokens by 62% and accuracy by 8.3 points.
- for who
- For anyone tracking frontier model reliability and post-launch performance.
- why now
- With Opus 5.5 released on September 22, this benchmark offers day-zero nerf detection.
- what changes
- They get a pre-registered day-0 baseline to measure performance drift, replacing speculation with statistical evidence.
- to do
- Monitor the daily results on the livenerf repo or run the benchmark to check for model drift.
key points
- livenerf benchmark uses 78 calibrated GPQA and math questions to detect post-launch drift
- Detects 7.5-point accuracy change per 10-day window; lower effort cuts tokens 62%
- Pre-registered protocol, runs daily for 30 days on Claude Opus 5.5
#model drift detection#anthropic#benchmark
score
score 8 out of 10. 0-10: how dense the facts are, multiplied by how much you can do with them after reading. 8+ means the topic's evidence bar is met: benchmarks and availability for a new model, amount and investors for a funding round, revenue figures for a solo-money story. Below 5 an item does not enter the digest. A press release scores 3 or less, a reprint loses 2, anything older than 14 days loses 1, a headline that misleads loses 3.
read the source