You know how your phone's camera gets better every year, but the phone costs the same? If the Bureau of Labor Statistics just tracked phone prices, they'd say inflation is flat. If they tracked price-per-megapixel, they'd see deflation. The gap between those two measurements is exactly the fight this paper picks — except the product is AI inference, the quality improvement is roughly 7× faster than what current methods capture, and the stakes are whether national productivity statistics are measuring reality or a shadow on the wall. The committed claim: statistical agencies are missing 87% of the AI inference price decline because they use matched-model methods (tracking the same product over time) instead of quality-adjusted hedonic methods (tracking what you get per dollar). Louis Yiven Zhu assembles 21,024 posted-price observations across 3,208 models from 86 providers, links them to 4,605 benchmark scores via a latent quality index estimated from benchmark response patterns, and builds the first comprehensive quality-adjusted price index for the AI inference market. Matched-model: prices fell 0.10 log points per year. Quality-adjusted: 0.73 log points per year. That's not a rounding error — it's a 7× gap. But here's the knife twist that keeps this from being a simple good-news-for-buyers story. Counted per completed task — not per token, but per actual thing accomplished — the buyer's price stopped falling. Reasoning models (think chain-of-thought, multi-step inference) consume tokens far faster than token prices decline. The seller's price per token drops; the buyer's price per task flatlines or rises. This is the seller-buyer price divergence, and it's the most economically consequential finding in the paper. It means the productivity gains from cheaper AI may be significantly overstated if you measure at the API level instead of the task level. The architecture here is applied econometrics, not ML. Zhu borrows the hedonic price index framework that statistical agencies use for computers and software — the tradition running from Griliches through Triplett — but replaces the usual product-characteristic approach (GHz, RAM, cores) with a latent quality index estimated from benchmark evaluation patterns. This is a meaningful methodological choice: instead of treating benchmarks as features in a regression, he treats them as noisy signals of an underlying quality dimension, estimated via item response theory. The quality ladder is built from evaluations, not specs. The integrity story is unusually strong for an economics preprint. The validity audit was pre-registered. The sharpest finding comes from that audit: excluding contamination-flagged benchmarks leaves model rankings intact at 0.998 correlation — the leaderboard barely moves — yet shifts the price index by 0.49 log points per year. This is devastating for the standard defense of benchmark-based statistics. The AI evaluation community's go-to argument ('rankings are stable, so contamination doesn't matter') is shown to be irrelevant to economic measurement. Leaderboard stability and price-index validity are different questions, and contamination poisons the second while leaving the first untouched. All data, code, and provenance are archived on Zenodo and reproducible from public sources at zero cost. The obvious next experiment Zhu didn't run: transaction-level pricing. Everything here is posted prices — the sticker on the window, not what people actually pay. Volume discounts, enterprise contracts, commitment-use pricing, and the massive free tiers offered by labs to gain market share are all invisible. My honest read: this isn't a data limitation he could easily fix — transaction prices are proprietary, and no provider publishes them. He'd need partnerships with cloud providers or API aggregators to get at actual spend per task. That's a (c) — saving it for the next paper, or waiting for the data to become available. For the AI industry, the seller-buyer price divergence is the number to track. If reasoning-model token consumption keeps outpacing token price declines, then the narrative of 'AI keeps getting cheaper' is true for providers and false for buyers. That has direct implications for enterprise AI budgets, for how we measure AI's contribution to GDP, and for whether the current wave of AI investment is generating the productivity returns that justify it.