Imagine you're a photographer who shoots 500 bracketed exposures of every scene, then picks the best one. Your assistant watches you for a year and learns to nail the shot in 4 frames. That's distributional distillation for diffusion models: a student network learns to reproduce the teacher's output distribution in far fewer steps, not by imitating individual samples but by matching the entire statistical shape of the output. The committed claim here is specific and testable: for continuous-space diffusion language models (where tokens live in a continuous simplex rather than a discrete vocabulary), you can train a student model to match the teacher's reverse-KL divergence using two distinct gradient estimation strategies — one continuous (Simplex-DMD, using pathwise gradients through a softmax relaxation) and one discrete (Reinforce-DMD, using REINFORCE with a learned density ratio as baseline). Both share the same student architecture and objective. This is not a new model class; it's a compression technique for an existing one, but the dual-path formulation and the scale of the demonstrated gains are new. The ladder matters. On OpenWebText at 1,024 tokens, Simplex-DMD reaches generative perplexity 45.6 at unigram entropy 5.44 nats in just 4 NFEs — a 49% reduction versus the strongest evaluated diffusion baseline at matched entropy and sampling budget. At the higher-compute end, Reinforce-DMD hits perplexity 14.9 at entropy 5.00 nats with 256 NFEs, a 20% improvement. These are real numbers against named baselines under a controlled comparison protocol (matched entropy, matched NFEs). The honesty here is that autoregressive models still dominate on raw perplexity — the comparison is within the diffusion-LM family, not against GPT-class models. Architecturally, this sits squarely in the distribution-matching distillation family (DMD), adapted from the image domain to language. The key structural choice is the output parameterization: the student produces either continuous simplex vectors (soft token probabilities) or categorical samples, and this choice cascades through the entire gradient estimation pipeline. Simplex-DMD gets low-variance gradients through reparameterization but lives in a relaxed space; Reinforce-DMD samples discretely but needs variance reduction via a learned density ratio. The paper makes this tradeoff explicit and explores it thoroughly across step counts. Integrity is reasonable but not bulletproof. The evaluation uses OpenWebText with standard perplexity and entropy metrics — not a cherry-picked benchmark. The comparison protocol (matching entropy and NFEs) is honest and addresses a real methodological gap in prior work where papers compared at mismatched operating points. However, this is self-evaluation against the team's own diffusion baselines, not an independent replication. Code availability is not explicitly confirmed in the abstract. The entropy-matching protocol is itself a contribution, but it also means the numbers are only comparable within this specific framework. The milestone question is where this gets interesting. The gap between 4-NFE diffusion LMs (perplexity ~45) and autoregressive models (perplexity ~15-20 on similar data) is still enormous. The path to practical deployment requires closing that gap while maintaining the parallel-generation advantage. The next meaningful threshold would be getting 4-NFE generation below perplexity 25 on OpenWebText, which would make the speed-quality tradeoff competitive enough for real applications like speculative decoding or draft generation. The obvious experiment not run: applying this to larger-scale diffusion LMs trained on more data, and directly benchmarking wall-clock generation speed against autoregressive models of comparable quality. The honest read is (a) — compute and access constraints. Training diffusion LMs at GPT-2/3 scale is expensive, and the base models this distills from are themselves research-scale. The gap between 'works on OpenWebText' and 'works at production scale' is where most distillation methods quietly die.