Imagine you're restoring a blurry photograph. You wouldn't start pixel by pixel — you'd first block in the big shapes (sky, ground, face), then progressively sharpen details within each region. This paper does exactly that for text generation: it runs two parallel denoising processes, one at the coarse "cluster" level (big shapes) and one at the fine "token" level (pixel detail), and lets them talk to each other throughout. The result is a continuous diffusion language model that finally closes a meaningful chunk of the gap to autoregressive and discrete-diffusion competitors. The committed claim: jointly diffusing a coarse semantic representation alongside fine-grained tokens, with per-modality schedules and samplers, yields large empirical gains in continuous diffusion language modeling — without significant compute overhead. Applied to CoBit as H-CoBit, the method drops generative perplexity from ~73 to 49.4 on LM1B and from ~71 to 50.4 on OpenWebText, improvements of 24.2 and 20.7 points respectively. On GSM8K math reasoning, H-CoBit hits 27.4% accuracy, surpassing all prior continuous diffusion and flow-matching models. The framework also generalizes: applied to FLM (a flow-matching model), H-FLM shows consistent gains, demonstrating this isn't a one-trick pony tied to a single architecture. The architectural insight is elegant in its simplicity. The authors cluster pretrained token embeddings (using k-means on e.g. RoBERTa embeddings) to create a coarser vocabulary — think of it as a semantic zoom-out where "dog," "puppy," and "hound" might share one cluster centroid. They then run continuous diffusion jointly over both the original token embeddings and these cluster embeddings, using a shared transformer backbone with modality-specific output heads. The key flexibility: each modality gets its own noise schedule and sampler, so the coarse signal can resolve earlier and guide the fine-grained tokens. This is a lightweight wrapper — parameter overhead is minimal, and the shared backbone means you're not doubling compute. On the ladder: continuous DLMs have historically trailed discrete diffusion models (like MDLM, UDLM) and especially autoregressive models. H-CoBit doesn't just beat the CoBit baseline — it surpasses discrete DLMs of comparable size on MAUVE and GenPPL. The 27.4% GSM8K accuracy outperforms SEDD (21.8%) and other continuous/flow competitors. That said, autoregressive models at the same parameter count still dominate on reasoning benchmarks, and the paper is honest about this gap. The real news is that continuous models are no longer embarrassingly behind. Integrity is solid but not watertight. The authors benchmark on community-standard datasets (LM1B, OWT, GSM8K, Lambada), compare against recent strong baselines (CoBit, FLM, MDLM, SEDD), and report multiple metrics including MAUVE, GenPPL, and zero-shot accuracy. Code is promised on GitHub. However, this is self-evaluated — no independent replication — and the cluster-count hyperparameter (number of k-means clusters) introduces a tuning dimension that could invite selection bias. The paper does ablate over cluster counts and schedules, which mitigates but doesn't eliminate this concern. The milestone question is where to look next. Continuous DLMs need to demonstrate scaling: can H-CDLM's gains persist at 1B+ parameters and on harder benchmarks (MMLU, HumanEval, long-form generation)? The current experiments are at ~170M parameter scale. If the hierarchical diffusion trick holds at 1B+ parameters with proportional GenPPL improvements, it becomes a serious architectural pattern. If gains plateau at scale, it's a nice trick for small models. The authors didn't run the scaling experiment — most likely compute-limited, given this is a two-author academic paper — and that's the single most important open question.