Imagine you're baking bread with a recipe that says "knead for 60 minutes." You've been buying expensive new flour brands, switching ovens, and redesigning your kitchen — all to make the bread rise faster. Then someone points out that your oven thermostat was miscalibrated the whole time. Turn the dial correctly and the original flour works fine in 10 minutes. That's the core mechanism here: the diffusion language model (DLM) community has been racing to build better models for few-step generation, when the real bottleneck was a suboptimal sampler temperature setting. The committed claim: a years-old masked diffusion language model (MDLM), with nothing but sampler sharpening — no retraining, no distillation, no architectural changes — achieves lower generative perplexity in 16 steps than its own standard sampler gets in 1024 steps. That's a 64× step reduction from a configuration change. The paper further demonstrates that this tuned baseline rivals or beats supposedly improved successor models like MDLM distilled variants and newer DLMs that required significant additional training compute. The architectural family is masked diffusion models — discrete denoising processes that iteratively unmask tokens in parallel, sitting in the broader family of non-autoregressive text generation. The key insight is mechanistic: when you unmask multiple tokens simultaneously, you destroy conditional dependencies between them. The resulting joint distribution is flatter than the true data distribution. Temperature=1 (the field's default) is therefore provably suboptimal even for a perfect denoiser. Sharpening (temperature < 1) compensates for the dependency destruction. The authors prove this theoretically for the case of parallel sampling from an exact denoiser. The paper's most provocative contribution may be methodological rather than empirical. The authors demonstrate that conventional per-output metrics (perplexity, MAUVE, entropy) can be gamed by any generator supported on as few as two outputs — meaning a degenerate generator can appear to optimally trade off quality and diversity. Their proposed alternative, GroupEval, separately measures within-output quality and across-output semantic diversity, revealing that a distilled model achieving 1.5–4.7× perplexity improvements shows no corresponding quality gain under human-aligned evaluation. This is a direct challenge to how the DLM subfield has been keeping score. Integrity is mixed. The theoretical result (suboptimality of temperature=1) is proved mathematically, which is strong. The empirical evaluation uses standard benchmarks and compares against named baselines including MDLM, its distilled variants, and other DLMs. However, the work comes from the same ecosystem as the original MDLM, validation is internal, code availability and pre-registration status are not stated in the abstract, and the GroupEval metric is introduced by the same authors who benefit from its use — a classic grading-your-own-homework concern. The field fight here is real: should the DLM community invest in better models or better samplers? This paper argues forcefully for the latter, at least until sampler optimization is exhausted. The implication is that a significant fraction of recent DLM papers may have been solving a problem that didn't exist — attributing sampler misconfiguration to model inadequacy. If this holds up under independent replication, it reshuffles the priority stack for the entire subfield. The successor experiment the authors didn't run is the obvious one: apply the same sampler tuning protocol to autoregressive baselines and large-scale DLMs (not just MDLM) to see if the gap closes universally or if MDLM is a special case. The honest read is probably (a) — compute and scope constraints — since systematically tuning samplers across many model families is expensive and this paper already carries a theoretical proof, a new evaluation framework, and extensive empirical results. But until someone does that cross-family comparison, the generality of the claim is uncertain.