21 signals
HOlO V1 IS LIVEone ranked AI digest a day, scored in publicREAD HOW IT WORKS →HOlO V2 STARTSyour X account, your signals, every day
signalHacker News Show HN2026-09-28

Show HN: Has Anthropic been nerfing their models without disclosure

A new deterministic benchmark, livenerf, tracks whether Anthropic's Claude Opus 5.5 degrades after launch. It runs daily for 30 days, using a panel of 78 calibrated questions, and can detect an accuracy change of about 7.5 points per 10-day window. Early data shows lower effort reduces output tokens by 62% and accuracy by 8.3 points.

for who
For anyone tracking frontier model reliability and post-launch performance.
why now
With Opus 5.5 released on September 22, this benchmark offers day-zero nerf detection.
what changes
They get a pre-registered day-0 baseline to measure performance drift, replacing speculation with statistical evidence.
to do
Monitor the daily results on the livenerf repo or run the benchmark to check for model drift.
key points
  • livenerf benchmark uses 78 calibrated GPQA and math questions to detect post-launch drift
  • Detects 7.5-point accuracy change per 10-day window; lower effort cuts tokens 62%
  • Pre-registered protocol, runs daily for 30 days on Claude Opus 5.5
#model drift detection#anthropic#benchmark
score
score 8 out of 10. 0-10: how dense the facts are, multiplied by how much you can do with them after reading. 8+ means the topic's evidence bar is met: benchmarks and availability for a new model, amount and investors for a funding round, revenue figures for a solo-money story. Below 5 an item does not enter the digest. A press release scores 3 or less, a reprint loses 2, anything older than 14 days loses 1, a headline that misleads loses 3.
read the source