Imagine you're running a city's water system, and every pipe has a slightly leaky valve. Most leaks evaporate quickly — a drip at the end of a dead-end street barely matters. But a leak at the main trunk line compounds: every downstream neighborhood gets progressively less pressure. STEPQuant is the insight that you don't need to replace every valve with an expensive brass one — you just need to figure out which pipes are trunk lines (spatial impact) and which leaks persist for miles (temporal persistence), then spend your repair budget there. The committed claim: post-training quantization of delta-rule linear attention recurrent states can match FP32 accuracy at a nominal 6-bit budget and beat uniform INT8 at 4 bits, by jointly reasoning about where errors propagate temporally (long-lived memory rows accumulate rounding damage across many decoding steps) and spatially (different key rows and value columns have wildly different magnitudes and downstream impact on model outputs). This is not the first paper to quantize recurrent states, but it is the first to decompose the quantization error landscape into these two complementary dimensions and allocate precision accordingly. The architectural family here is delta-rule linear attention — the branch of sub-quadratic sequence models (relatives of Mamba, RWKV, and RetNet) that maintain a fixed-size recurrent state matrix instead of a growing KV cache. The key hardware property being exploited is that this state matrix, while fixed-size per layer, becomes the dominant memory bottleneck under concurrent serving when batch sizes grow. STEPQuant's mechanism fits key-row and value-column quantization scales jointly, calibrated on state distributions and a sensitivity metric for each key row's contribution to output error. It's post-training — no retraining required — and ships with optimized CUDA kernels integrated into the SGLang serving framework. The ladder is honest and well-constructed. They test on two production-scale models: Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct. The headline result is that 6-bit STEPQuant matches FP32 states across both long- and short-generation benchmarks, while their 4-bit configuration outperforms naive uniform INT8 quantization. The baselines include the obvious ones: FP32 full-precision states and uniform INT8. The gap between 'matches FP32 at 6 bits' and 'beats INT8 at 4 bits' is the real contribution — it means you can halve the bit budget below INT8 and still come out ahead because you're spending bits where they matter. Integrity is solid for a technical report. The experiments span multiple model scales, both long-context and short-generation benchmarks, and the code is publicly released on GitHub. The main caveat is that this is self-validation — no independent replication yet — and the benchmark selection may favor the method's strengths. The temporal-spatial decomposition is well-motivated empirically but the discovery of which rows matter most is calibration-dependent, and we don't know how sensitive the results are to calibration set choice. The practical milestone is clear: 5× recurrent-state compression with up to 68.7% total serving memory reduction. The next number to watch is whether this extends cleanly to 3-bit or 2-bit configurations — that's the threshold where linear attention models could serve at costs competitive with heavily optimized KV-cache quantization for standard Transformers, potentially changing the cost calculus for which architecture you deploy. The gap is probably 6-12 months of kernel optimization and mixed-precision engineering. The obvious experiment not run: applying STEPQuant to other linear attention families beyond delta-rule (e.g., pure Mamba-style SSMs, RetNet, RWKV). The temporal-spatial decomposition should transfer in principle, but the sensitivity structure may differ. My read: they're saving this for the next paper. The framework is general enough that extending it is straightforward engineering, and demonstrating breadth across architectures is a clean follow-up publication.