Imagine you're assembling a jigsaw puzzle in a dark room where every time you pick up a piece, you have to add random noise to your sense of touch — and once you place a piece, you can never move it. That's DP-OPT: greedy, sequential, and one bad noisy placement cascades into a ruined picture. Now imagine instead you have ten friends each assembling their own full puzzle in well-lit rooms, and the only thing that happens in the dark is you briefly glance at the finished puzzles to rank them. That's DP-ES — the privacy noise hits the scoring, not the construction. The committed claim: under a conservative differential privacy budget of ε≤1.0, DP-ES achieves 88.1% on GSM8K math reasoning (versus DP-OPT's wildly unstable 49.5±28.5%), while also hitting 99.7% on MedQA, 73.5% on BANKING77, and 86.8% on Alpaca. The standard deviation collapses by roughly 9×. This is not a marginal improvement — it is a structural fix for a known failure mode. The diagnosis of WHY DP-OPT fails is as valuable as the replacement. The authors identify two specific pathologies: prompt-template drift (the greedy construction wanders off-template as noisy token selections accumulate) and noise-sensitive irreversible choices (early token decisions lock in bad directions with no recovery path). These are not implementation bugs — they are structural consequences of applying DP noise to a sequential, greedy search. The paper logged a full DP-OPT search trajectory to demonstrate this, which is more diagnostic honesty than most papers in this space offer. Architecturally, DP-ES belongs to the evolution strategies family — specifically zeroth-order black-box optimization. The key structural insight is a clean separation of concerns: mutation happens via LLM calls that never touch private data (so mutation is free from a privacy standpoint), and the ONLY privacy expenditure is on Gaussian-noise-protected evaluation scores. Selection is deterministic or Gumbel-smoothed, which qualifies as post-processing under DP composition rules. This means the privacy accounting is dramatically simpler and tighter than DP-OPT's token-by-token mechanism. The ladder position is strong against the named baseline. DP-OPT is the current standard for DP prompt optimization, and DP-ES beats it by 38.6 percentage points on GSM8K with 9× lower variance, while being 2.5× faster and using 3.3× fewer private-data call groups. The paper is honest that benchmarks like MedQA at 99.7% may be near saturation, which limits signal. What's missing is comparison against non-DP prompt optimization methods — we don't know how much accuracy DP-ES sacrifices for privacy versus the unprotected baseline. Integrity checks are more thorough than typical. The authors run selection ablations, population size ablations, implementation-level noise verification, and a 200-profile exact-match memorization stress test to confirm the formal DP guarantee holds in practice. The scope limitation is stated clearly: all experiments use public benchmarks as proxies for private data. The paper explicitly flags that end-to-end validation on genuinely sensitive, non-saturated deployment data is future work. This is the right thing to say and also the biggest hole. The obvious next experiment is running DP-ES on actual sensitive deployment data — medical records, financial transactions, user queries — where the privacy guarantee is load-bearing rather than academic. The authors didn't run this, and the honest read is (a): they don't have access to such data in a research setting where they could publish results, and the IRB/compliance overhead for genuinely private datasets is substantial. This is a real barrier, not a dodge. The second missing experiment is scaling to harder, non-saturated benchmarks where the accuracy gap between DP and non-DP methods would be more visible.