Imagine you're building a word search puzzle. You have a grid, and you need to fill it in one cell at a time so that every row and column spells a valid word. At each step, you pick a cell and fill it with a letter that's locally legal — it doesn't break any row or column constraint right now. But when the grid is done, you notice the distribution of completed puzzles is wrong: some valid grids show up way too often and others almost never, even though your per-step choices were all 'correct.' The problem isn't any single step — it's that each step silently narrows future options in ways that accumulate into a systematic skew. That's the trajectory bias this paper identifies and fixes in masked diffusion language models. The committed claim: existing 'step-exact' constrained decoders for Masked Diffusion Language Models (MDLMs) — methods that guarantee each individual unmasking step satisfies the constraint — produce trajectories whose overall distribution is biased away from the model's true conditional distribution over valid outputs. The authors prove this bias exists, derive an exact closed-form expression for it as a product of ratios measuring how valid-continuation mass shifts when the denoiser is reconditioned, and characterize the precise conditions under which the bias vanishes. The fix is TWISTER, the first automaton-twisted Sequential Monte Carlo (SMC) decoder for MDLMs. The key architectural insight: the step-exact decoder becomes the proposal distribution inside an SMC sampler, and the Feynman-Kac correction weights — which would normally be intractable — turn out to be exactly computable for regular language constraints using quantities already computed during step-exact sampling. This is not a hack or approximation; the authors prove the resulting Feynman-Kac model targets the unbiased Doob h-transformed path law conditioned on constraint satisfaction. The method lives in the family of sequential Monte Carlo methods applied to discrete diffusion, leaning on the factorized mean-field posterior structure of MDLMs and the chain-structured factor graph that automaton constraints provide. The integrity story is strong in one dimension and weaker in another. The theoretical contribution — proving trajectory bias exists and deriving its exact form — is a mathematical result, not an empirical claim vulnerable to benchmark shopping. The correction's exactness is also proven, not merely demonstrated. However, the abstract provides no empirical benchmarks, no dataset names, no generation quality numbers, and no comparison against specific prior methods by name with metrics. The paper is under review, which explains the absence of community vetting, but makes it impossible to assess practical performance claims. The ladder position is interesting because it doesn't compete on generation quality numbers — it competes on correctness of the sampling procedure itself. The prior art is the step-exact constrained decoding line of work for MDLMs (which enforces constraints via automaton-structured dynamic programming at each step). TWISTER doesn't claim better perplexity or BLEU; it claims the existing method is solving the wrong problem and then solves the right one. This is a foundational correction, not an incremental benchmark improvement. The milestone to watch is whether TWISTER's SMC correction changes downstream generation quality in practice. The theoretical bias is proven to exist, but its practical magnitude across real tasks — code generation, structured text, formal language synthesis — is the open question. If the bias is small for most practical constraints, the contribution remains theoretically clean but practically marginal. If the bias is large for constraints practitioners actually care about (e.g., syntactically valid code, schema-compliant JSON), this becomes load-bearing infrastructure for constrained generation in diffusion models. The obvious experiment not run: empirical evaluation showing the magnitude of trajectory bias on real constrained generation tasks and demonstrating that TWISTER measurably improves sample quality or distribution fidelity compared to step-exact baselines. The honest read is (c) — this is a theory-first paper under review, and the empirical evaluation is either in supplementary material not included in the abstract or being saved for the camera-ready version. The absence of any named benchmark or metric in the abstract is the tell.