Imagine you train a drug-sniffing dog on cocaine and heroin. It's great — 99% hit rate. Then a new synthetic opioid hits the street and nobody retrains the dog. Suddenly it walks past most contraband like it's air freshener. That's not the dog failing; that's the training set going stale. This paper is about exactly that dynamic, except the dog is an AI-text detector and the new drug is the next ChatGPT release. The committed claim: LLM-text detectors, calibrated at a 1% false-positive rate on human text, can collapse from above 99% true-positive detection to 3.8% when a new model generation appears — not because the detector is bad, but because model turnover renders its training obsolete. This isn't about adversarial attacks or deliberate evasion. It's plain version drift doing the damage. Nakajima and Mizuno built a controlled testbed: 4,000 pre-ChatGPT abstracts from PNAS, each rewritten by 23 LLM versions spanning three vendors (June 2023 to August 2026). They then trained classifiers under several maintenance regimes — from the ideal (retrain on every new version) to the realistic (train once, never update). The sharpest finding is at generational boundaries: a detector calibrated on GPT-4-class models catches almost nothing from the first GPT-4o-class model. Going backward is also broken — detectors trained on newer versions miss older rewrites. Vocabulary fingerprints explain most of the transfer patterns. The practical scenario simulations are where the paper bites hardest. When a single detector is asked to cover all 23 versions simultaneously, you get an impossible tradeoff: either flag 12.5% of legitimate human-written abstracts (roughly one in eight), or miss 33% of rewrites from the newest model. A commercial detector they tested confirmed the pattern — it missed most rewrites from the version just after the sharpest generational boundary while barely flagging any human text. The screening looked fine on its own benchmarks and was quietly blind to the newest threat. The architecture is straightforward: supervised binary classifiers (human vs. LLM-rewritten) operating on vocabulary-level features. The power of the paper isn't in the classifier design — it's in the experimental design around version coverage. By systematically varying which versions appear in the training set, they isolate the version-turnover effect from detector quality. This is a measurement paper, not a methods paper. Integrity is solid for the claim being made. The 4,000 PNAS abstracts are a real, pre-ChatGPT corpus. The 23 LLM versions are named and span three vendors. The maintenance scenarios cover the realistic range. The one limitation is that this is all abstracts — short, structured, domain-specific text. Detection on full papers, essays, or creative writing may behave differently. The authors don't overreach. The policy implication is blunt: any journal or conference treating a detector's benchmark accuracy as durable is operating on a false premise. Benchmark accuracy is a snapshot, not a property. Every new LLM release invalidates it, and the invalidation can be catastrophic, not gradual. The paper recommends re-verification with every release — including re-testing against older versions, since backward compatibility also breaks.