Imagine you're training a basketball player by having them practice free throws blindfolded, with the ball placed randomly in their hands each time — but then expecting them to drain shots in a real game where they dribbled to the line themselves, eyes open, remembering every step. That's the core absurdity masked diffusion language models (MDMs) have lived with: trained on randomly masked sequences, deployed on trajectories shaped by their own predictions. PUMBA is the coaching regime that finally puts them in scrimmage. The committed claim: by training the denoiser on consecutive steps of its own policy-induced trajectories — passing continuous information between steps and backpropagating through time (BPTT) — you can close the train-inference mismatch that has kept MDMs behind autoregressive models. This isn't the first paper to notice the gap; MDLM, GenTron, and others have poked at pieces of it. PUMBA's contribution is unifying three levers (trajectory alignment, inter-step information, and multi-step optimization) into a single controlled study that isolates what actually matters. The findings are specific and non-obvious. Exact train-inference alignment — where you literally sample from the model's own policy during training — fails due to local overfitting. A looser alignment strategy works better by bringing training masks closer to inference masks without collapsing onto them. Continuous information passing between denoising steps beats discrete gradient estimators (Gumbel-Softmax, straight-through). And performance scales with the number of BPTT steps, which they back with a theoretical argument about variance reduction. On the ladder: at the 110M-parameter scale, PUMBA matches the best checkpoint of a same-size autoregressive model — a meaningful bar for MDMs, which have historically trailed AR models on perplexity and downstream quality. Scaled to LLaDA-8B via supervised fine-tuning, it delivers up to 22% fewer neural function evaluations (NFEs) at matched performance in full-canvas generation and 26% fewer in block diffusion. That's not a quality leap; it's an efficiency leap at parity, which matters when inference cost is the bottleneck. Architecturally, PUMBA lives in the masked discrete diffusion family — specifically the absorbing-state variant where tokens transition to a [MASK] state and denoising reverses the process. The key structural choice is BPTT across denoising steps, which is expensive (memory scales linearly with trajectory length) but lets gradients flow through the sequential decision chain rather than treating each step independently. The inter-step signal is continuous logits, not argmax tokens — a choice that avoids the well-known gradient estimation headache of discrete latent variables. Integrity is decent for a methods paper. The controlled ablation at 110M parameters is thorough: they isolate each of the three components and report interactions. The LLaDA-8B scaling experiment uses established benchmarks (MMLU, GSM8K, HumanEval) and compares against the publicly available LLaDA baseline with matched and doubled training budgets. The main vulnerability is that all baselines are internal or from the LLaDA team — no independent replication yet, and the AR comparison at 110M is against their own trained model rather than a community checkpoint. The milestone to track: MDMs matching AR models at scale isn't just about perplexity — it's about the NFE-quality Pareto frontier. PUMBA gets 22-26% NFE savings at 8B. The next concrete threshold is whether this compounding holds at 70B+ scale with harder generation tasks (long-form reasoning, code), which would make parallel decoding competitive for real deployment. The obvious experiment they didn't run is PUMBA on a model larger than 8B or on reinforcement learning from human feedback (RLHF) — likely a compute budget constraint given that BPTT across multiple denoising steps at 8B is already expensive, but also possibly being saved for a follow-up that positions PUMBA as an alignment technique.