Imagine you're sculpting a face from clay. You could work by placing individual tiles in a mosaic — each tile chosen independently, hoping the whole face emerges (that's discrete diffusion). Or you could smooth the entire surface at once, but you can't see what face you're making until you step back at the very end (that's continuous diffusion). HC-DLM does something different: you sculpt the smooth surface, but at every step you hold up a mirror that shows you a mosaic approximation of your current work, and that approximation guides your next set of smoothing strokes. The latent is the clay; the tokens are the mirror. The committed claim: HC-DLM is the first diffusion language model that makes continuous latent space the sole persistent generative state while feeding discrete token readouts back as scaffolding at every denoising step, with a principled variational bound on token likelihood. This is not a minor architectural tweak. Prior continuous diffusion LMs (like CDCD or SEDD's continuous variants) denoise a shared state but remain blind to valid token configurations until final decoding. Prior discrete diffusion models (MDLM, SEDD) sample tokens independently in parallel, severing inter-token dependencies. HC-DLM's hierarchy — continuous state generates tokens, tokens feed back into continuous state — breaks that dichotomy. The ladder matters here. On Sudoku (9×9 board completion), the paper reports puzzle-level accuracy improvements over discrete baselines like MDLM. On Countdown (a mathematical planning task requiring multi-step arithmetic), HC-DLM again outperforms matched-size discrete and continuous diffusion models. On LM1B, a standard language modeling benchmark, it improves generative perplexity over baselines at matched model size. The abstract doesn't give exact delta numbers, which is a flag — we're told 'improves over' but not by how much. The project page may contain specifics, but the abstract leaves the reader estimating rather than measuring. Architecturally, this is a variational continuous diffusion model with a discrete readout head. It belongs to the score-based / denoising diffusion family, but with a hierarchical twist: the denoiser operates in continuous space (likely a transformer backbone processing continuous embeddings), while a token-projection layer maps the latent to discrete token distributions at each step. The training objective derives from a variational bound — specifically, a bound on log-likelihood of the token sequence — which gives it principled training rather than the heuristic losses some continuous diffusion LMs rely on. The key structural bet is making the latent the ONLY persistent state, with tokens as ephemeral scaffolding rather than the generative target itself. Integrity is mixed. The three evaluation domains — Sudoku, Countdown, LM1B — span structured reasoning and open-domain language, which is good breadth. Sudoku and Countdown are deterministic-answer tasks where success is unambiguous, reducing cherry-picking risk. LM1B is a community benchmark. But the abstract doesn't name exact baselines with version numbers or report numeric deltas, and there's no mention of code release (though a project page exists). The 'matched model size' framing is important and honest — it controls for the obvious confounder — but we'd want to see parameter counts, training compute, and step counts to verify the match is real. The milestone question for diffusion language models is whether they can close the gap with autoregressive models on open-ended generation quality while maintaining their parallel-decoding and bidirectional-reasoning advantages. HC-DLM's LM1B result is a step, but the real unlock would be competitive perplexity on larger-scale benchmarks (C4, The Pile) at GPT-2-scale or beyond, with wall-clock speedups from parallel decoding. The field is probably 2-3 years from diffusion LMs being a practical alternative to autoregressive generation for production use. The obvious experiment not run: scaling HC-DLM to larger model sizes and longer sequences. Sudoku is 81 tokens. Countdown is short. LM1B sentences are typically under 30 tokens. The question every reader should ask is whether the hierarchical feedback loop — tokens feeding back into continuous state at every step — scales gracefully or becomes a bottleneck at 512+ token sequences. The honest read: likely (a) compute-limited at this stage, but possibly (c) the scaling story is the next paper, since the architectural contribution stands on its own at smaller scale.