For months, a recurring suspicion has circulated in AI communities: that frontier models get quietly worse after launch. Quantization, routing changes, smaller models behind the same API name — the theories vary, but the common thread is that nobody has had a clean day-zero baseline to test against. Every argument devolves into anecdotes. LiveNerf is a small open-source project designed to fix exactly that, starting with Claude Opus 5.5's release on 2026-09-22. The approach is methodologically disciplined. From a pool of 2,336 questions drawn from GPQA Diamond, MMLU-Pro, competition math, and AIME 2025-26, the benchmark screened down to 78 questions that Opus 5.5 gets right only sometimes — the informative ones. Questions the model always aces or always flunks carry zero signal about drift, so they're excluded. The panel was locked, confirmed, and validated under a pre-registered protocol (v2) before the first run, and the public git timestamp serves as the pre-registration record. The design choices are deliberately boring and that is the point. Every run uses frozen prompts, a pinned CLI version (2.1.280), exact graders, and raw logs that are append-only and never edited. There is no LLM judge — scores are exact-match. The statistical framework follows Anthropic's own 'Adding Error Bars to Evals' paper, and the infrastructure is built on the UK AI Security Institute's open-source Inspect framework. The harness hash (461391b6fce64167) has been identical across all six days collected so far. The instrument has known, honestly reported limits. It can detect an accuracy change of roughly 7.5 points per 10-day window. A same-family model swap — Opus 5 substituted for Opus 5.5 — was not distinguishable at 99% confidence in validation samples, showing only -3.8 ± 6.3 points and -23% fewer tokens. The more sensitive signal turns out to be output token count: lowering effort from high to low cuts output tokens by 62% while only dropping accuracy by 8.3 ± 4.5 points. If a provider starts serving a cheaper model, token count will likely move before accuracy does. The benchmark also surfaces uncomfortable truths about its own question pool. An audit of the 78-question panel found 8 answer keys that appear wrong and 30 ambiguous questions — expected when you're selecting for items a strong model only sometimes gets 'right.' Rather than quietly dropping them, the project keeps all questions in and pre-registers a sensitivity analysis to rerun results without the suspect items. The broader significance is structural. This is an accountability tool for an industry where model providers control the serving infrastructure and users have no contractual guarantee that the model behind an API name stays the same. LiveNerf doesn't claim to solve this — it claims to make one specific kind of degradation statistically detectable, and to do so with enough methodological rigor that the results are hard to dismiss. The first possible call arrives around 2026-10-24, after 20 days of data. As of September 29, 2026, six of 30 planned days have been collected, none missed, all running the full 90 samples per day. The series runs once daily for 30 days: days 1-10 form the baseline, then two 10-day comparison windows follow. The project runs on a Claude Max subscription through headless Claude Code, with no API key — meaning it measures exactly what a paying consumer gets, not a privileged research endpoint.