Imagine you're running a restaurant kitchen where a prep cook guesses what the head chef will need next. If the prep cook guesses right, the whole line moves faster — dishes fly out. If the guess is wrong, the head chef corrects and the prep cook's wasted effort costs a beat. Now imagine training that prep cook. You could grade each individual guess in isolation, or you could grade the whole shift — how many rounds of correction did the kitchen need to finish the night's orders? The second approach is obviously better, but it requires tracking how each early mistake cascades into later rounds. That's the core problem this paper solves for speculative decoding. The committed claim: parallel draft models for speculative decoding have been trained with block-local surrogate objectives that ignore the coupling between decoding rounds. This paper derives the first exact training objective — Expected Decoding Rounds (EDR) — that directly minimizes the global number of verification rounds, with no auxiliary hyperparameters. It then shows an unbiased temporal-difference gradient that makes this objective practical to optimize. The theoretical machinery is clean. The authors model speculative decoding as a Markov reward process where states correspond to acceptance patterns within a block, transitions depend on the draft distribution, and rewards correspond to accepted tokens. The EDR objective falls out naturally as the expected number of rounds to decode a sequence, weighted by the stationary state occupancy. This is not a loose upper bound or a proxy — it's the thing you actually care about. The temporal-difference gradient they derive allows stochastic optimization from target-model rollouts, and the same Markov framework yields an exact offline evaluator that enables paired drafter comparisons without running speculative decoding end-to-end. Where it lands on the ladder: the authors finetune two state-of-the-art parallel drafters — DSpark and DFly — using EDR and compare against existing training objectives across nine benchmarks spanning math reasoning (GSM8K, MATH500), code generation (HumanEval, MBPP, MultiPL-E), and chat (MT-Bench, AlpacaEval). EDR-finetuned models consistently improve mean accepted length over both the original drafters and drafters finetuned with prior surrogate losses. The baselines here are real — DSpark and DFly are current, and the comparison objectives include the alternatives practitioners actually use. The integrity profile is solid but has the usual ML-paper gaps. Nine benchmarks is respectable coverage, the baselines are current and named, and the Markov reward process formulation provides mathematical grounding rather than just empirical claims. However, this is same-team evaluation with no independent replication, no pre-registration, and the code/model availability is not confirmed from the abstract alone. The offline evaluator they derive is actually a methodological contribution in its own right — it makes future comparisons cheaper and more controlled, which is a genuine service to the community. The milestone question is concrete: speculative decoding matters because it's the primary pathway to making large-model inference cheaper without distillation or quantization. The practical target is wall-clock speedup on real serving infrastructure. Current parallel drafters achieve roughly 2-3× speedup; the question is whether principled training pushes this toward 4-5× or whether the gains plateau as draft models get better. EDR's contribution is removing one bottleneck (the training objective mismatch), but hardware integration, serving-system optimization, and draft-model architecture are independent bottlenecks that remain. The obvious next experiment is scaling to significantly larger target models (70B+, 400B+) and measuring wall-clock improvement on production serving stacks rather than just mean accepted length. The authors likely stayed at manageable scale because EDR requires target-model rollouts during training, which gets expensive fast with very large targets. This is compute-budget-limited, not idea-limited — expect this to appear in a follow-up or in industry adoption.