Imagine you're a chef prepping for dinner service by taste-testing dishes made from yesterday's ingredients — then being surprised when tonight's fresh produce cooks differently. That's roughly what standard LLM pruning does: it decides which weights to zero out by measuring their importance on pre-collected human-written text (the calibration set), then deploys the pruned model to generate its own tokens, which have a different statistical distribution. SparseDecoding's core insight is embarrassingly obvious once stated: calibrate the pruner on the model's own autoregressive output, not on someone else's text. The committed claim is two-pronged. First, an algorithmic fix: collect layer-wise activations during dense-model autoregressive generation (excluding prefill) and use those to build the Hessian that guides pruning. This eliminates the distribution shift between calibration and deployment. Second, a systems fix: an optimized N:M sparse matrix-vector (SpMV) kernel using bitmask indexing and fixed-step traversal, designed specifically for the decoding bottleneck where you're multiplying a sparse matrix by a single vector, not a batch of vectors. Why does this matter? Decoding — generating one token at a time — is memory-bound, not compute-bound. Every parameter the GPU reads from memory costs latency. Pruning reduces the number of nonzero weights, but most existing sparse kernels are optimized for sparse matrix-matrix (SpMM) multiplication, which dominates the prefill stage. Decoding is dominated by SpMV, and until now the kernel support for actually exploiting sparsity in SpMV has been weak. SparseDecoding attacks both halves: better pruning decisions AND a kernel that can actually capitalize on the resulting sparsity during the operation that matters. The results span four models — Llama-3.1-8B, Llama-3.3-70B, Qwen3-14B, and Qwen3-32B — across long-form generation benchmarks. The decoding-aware calibration consistently outperforms standard fixed-text calibration on generation quality. The SpMV kernel delivers up to 1.48× end-to-end wall-clock decoding speedup on A100 GPUs. These are N:M structured sparsity patterns, meaning fixed ratios of zeros within small blocks, which hardware can exploit with predictable memory access patterns. The ladder here is honest but modest. The comparison is against standard Hessian-guided pruning (the SparseGPT / Wanda family) with conventional calibration — the current mainstream approach. SparseDecoding beats it on generation benchmarks, which is the right metric for a decoding-focused method. But the paper is careful to position itself as an improvement within the existing pruning paradigm, not a replacement for it. The 1.48× speedup is real and measured end-to-end, not just kernel-level, which is refreshingly honest — many pruning papers report theoretical FLOPs reductions that never materialize as wall-clock gains. The integrity profile is reasonable for a systems-meets-algorithms paper. Four representative models spanning 8B to 70B parameters, generation-quality benchmarks (not just perplexity on WikiText), and actual wall-clock measurements on real hardware. The main gap: no independent replication yet, and the benchmarks appear to be chosen by the authors rather than pre-registered. The distribution-shift observation is well-motivated but the magnitude of the effect could vary across tasks and model families in ways not fully explored. The obvious next experiment is scaling to even larger models (405B+) and testing on different hardware (H100, Blackwell) where the memory bandwidth characteristics differ. The authors likely ran out of compute for 405B experiments — that's a plausible and forgivable gap. The more interesting missing experiment is a systematic ablation showing exactly how much of the quality gain comes from decoding-aware calibration versus the SpMV kernel, across varying sparsity ratios and sequence lengths.