Imagine you're running a factory where one massive chalkboard tracks every order ever placed. Every time a new order arrives, a clerk reads the entire board, updates it, and writes it back. The board is the bottleneck — not the math, but the physical act of reading and writing all that chalk. LeapQuant's insight is: what if the clerk only rewrites the board every 64 orders, keeps a small notepad of recent changes in the meantime, and preserves the few most important entries in permanent ink so rounding errors from the rewrite don't compound? That's the paper. The committed claim: you can quantize the recurrent state of linear attention models (Gated DeltaNet, Kimi Delta Attention, and similar hybrid architectures) to 8-bit precision at inference time, with no retraining, and lose almost nothing in accuracy. The mechanism has two parts. Per-window quantization buffers a window of tokens (the notepad) in high precision, then quantizes the accumulated state only once at the window boundary — this amortizes the rounding error across many tokens instead of compounding it at every step. Compensator Tokens pull the biggest outlier rows and columns out of the state matrix and keep them in FP16/FP32, then smooth the residual before quantizing. Together, these two tricks keep quality near FP32 baseline across Qwen-3, Kimi-K2, and GLM-4 model families. The ladder here is straightforward: the baseline is FP32 inference of the same models, and the comparison metric is accuracy preservation plus wall-clock speedup. LeapQuant achieves 2.05–3.70× speedups at the kernel level and 1.47× end-to-end on three NVIDIA GPUs (B200, RTX PRO 6000, RTX 5090). The paper also compares against naïve per-token quantization and shows that approach collapses quality, validating that the window-based and compensator tricks are load-bearing. The honest limitation: end-to-end speedup is much lower than kernel speedup because memory bandwidth for the recurrent state is only part of the total pipeline. Architecturally, this sits squarely in the linear-attention / state-space model family — specifically the gated-linear-recurrence variants that compress KV context into a fixed-size state matrix. The method exploits a hardware property: modern GPUs are memory-bandwidth-bound when reading/writing large state tensors at every token, so shrinking the state from FP32 to INT8 directly reduces the bottleneck. The window buffering trades a small amount of compute (recomputing contributions from buffered tokens) for a large reduction in memory traffic. This is a systems optimization, not an architectural change. Integrity is reasonable for a systems paper. Experiments span three model families, multiple GPU architectures, and standard NLP benchmarks. The authors show both kernel-level and end-to-end numbers — the latter being the honest metric — and acknowledge the gap between them. There is no independent replication, which is expected at submission time. The benchmarks appear to be standard community tasks, not custom-designed ones. One concern: the window size (64, 128, etc.) is a hyperparameter whose optimal value could vary across models, and the paper does not fully explore sensitivity here. The milestone question for this subfield is not about LeapQuant specifically but about when linear-attention hybrids become the default serving architecture for long-context LLMs. LeapQuant's contribution is removing one obstacle — the memory cost of the recurrent state — but the broader question is whether linear attention quality matches standard attention at scale across all tasks. The 1.47× end-to-end speedup is real but not transformative; 2–3× end-to-end on real serving workloads would be the number that triggers production adoption. The obvious experiment not run: applying LeapQuant during training, not just inference. The authors explicitly scope this as training-free, which is pragmatic — training-time quantization is a much harder problem with gradient noise interactions — but also means we don't know whether further quality recovery is possible with quantization-aware fine-tuning. Honest read: this is (a) and (c) — training-time work is expensive and is almost certainly the next paper.