Imagine you're a talent scout at a music festival with 400 bands playing simultaneously. You can't watch every act, but you've grouped them into stages by genre. You watch one representative act per stage, then rush to the stages that looked promising to catch the full set. You weight your final ranking by how unlikely you were to visit each stage in the first place — correcting for the fact that you couldn't be everywhere. That's SANTA++, except the "festival" is a 32K-token KV cache and the "scout" is every attention query. The committed claim: you can reduce KV cache memory reads to 16-22% of dense attention's cost while retaining 94-99% of downstream task quality on long-context benchmarks, with no retraining, no architectural changes, and a mathematically grounded importance sampling correction that makes the approximation unbiased. This is not approximate attention through learned sparsity or low-rank projections — it's a sampling-theoretic approach applied at inference time on top of existing models. The method organizes cached keys into fixed "teams." Each team has a representative key. The query scores all representatives (cheap — there are far fewer representatives than total keys), then samples teams proportional to those scores. Within the sampled teams, exact attention is computed. The crucial trick is the inverse-probability reweighting: each team's contribution is scaled by the reciprocal of its inclusion probability, making the estimator unbiased in expectation. This is textbook importance sampling dressed in systems engineering clothing — GPU kernels, fused operations, memory-bandwidth awareness. The ladder is reasonably well-constructed. Dense FlashAttention is the baseline, which is the right one. On LongBench v2 and HELMET RAG, SANTA++ with 32-64 sampled teams retains 94-99% of dense scores using Qwen2.5-7B-Instruct at 32K context. RULER is harder: 85-91% retention. The 1.69× attention speedup at 32K context with 31 teams is a real systems number, not a theoretical FLOP count. What's missing: no comparison against other sparse or approximate attention methods (e.g., BigBird, Longformer, or recent KV cache compression techniques like StreamingLLM, H2O, or Scissorhands). The paper benchmarks against the densest possible baseline, which is honest but incomplete — the field already has sparse-attention methods. Integrity is mixed. The benchmarks — LongBench v2, HELMET, RULER — are community-standard long-context evaluation suites, which is good. But the evaluation is entirely self-reported on a single model family (Qwen2.5-7B-Instruct) at a single context length (32K). No independent replication. The kernels are released, which is a strong signal — reviewable code is the closest thing ML has to pre-registration. The RULER degradation (85-91% vs 94-99% on other benchmarks) is reported honestly, which builds trust. The milestone framing is implicit but extractable: the method reads 16-22% of KV cache at 32K context. The obvious next milestone is demonstrating the same quality retention at 128K or 1M context, where the memory-bandwidth bottleneck becomes genuinely punishing and the value proposition of reading fewer entries grows nonlinearly. The paper explicitly notes compatibility with compressed KV representations like multi-head latent attention — that composition experiment is the high-value next step. The experiment the authors didn't run: scaling to longer contexts (128K+) and combining SANTA++ with KV cache compression methods. The paper hints at this combination as "in principle" compatible but doesn't demonstrate it. My read: (a) compute/engineering budget — building fused GPU kernels is expensive, and demonstrating the composition requires integration work with other codebases. Possibly (c) — the MLA + SANTA++ composition is the obvious next paper.