Imagine you're lost in a pitch-black warehouse and need to find the lowest point on the floor. Backpropagation is like having a flashlight that shows you the exact slope under your feet — efficient, but you can only see what the light reveals. Evolution strategies are like hiring a thousand people to each take a random step and radio back whether they went downhill — powerful in principle, but you need a thousand people. Dust's trick is realizing that if you're already walking through the warehouse carrying a tray of a thousand marbles, you can nudge each marble independently and watch which ones roll downhill, all in one trip. Each marble is a token. One forward pass, thousands of gradient estimates. The committed claim: Dust is the first zeroth-order optimization method competitive with backpropagation at pretraining transformer language models. At large population sizes it even exceeds backprop in multiple settings. This is not a toy demonstration on MNIST — they pretrain actual transformers on language modeling, the hardest mainstream benchmark for credit assignment. The method perturbs activations (not weights) with Gaussian noise independently at every token position, scores each perturbation by its effect on that token's loss, and averages the reward-weighted perturbations to estimate gradients. The outer product of this estimated error with the layer input gives the weight gradient. Attention layers get special treatment: they're credited through estimated errors at the attention output over current and future tokens rather than through direct token loss. The ladder comparison is striking but requires careful reading. Dust is competitive with backprop — meaning it can match backprop's pretraining loss curves — but at substantially more compute. The paper is explicit about this: they do not claim compute efficiency today. Against weight-space evolution strategies like EGGROLL (Sarkar et al., 2025), Dust is 10³ to 10⁴ times more efficient from 1M tokens onward, based on their extrapolations. This is the right comparison class: Dust demolishes other zeroth-order methods but does not yet beat backprop on a FLOP-for-FLOP basis. The honest framing is 'first zeroth-order method in the game' rather than 'backprop replacement.' The most provocative finding challenges conventional wisdom about scaling. Zeroth-order methods are widely believed to fail as networks grow — the search space explodes and you drown in noise. Dust finds the opposite: a 243M-parameter model outperforms a 120× smaller model at most population sizes. Larger models are more population-efficient, not less. If this holds, it inverts the standard objection to search-based training and suggests overparameterization provides better loss-landscape geometry for zeroth-order search, not worse. The gradient estimates also align more closely with backprop's as population grows, and this alignment holds at every scale tested up to 1B tokens. Architecturally, Dust belongs to the node-perturbation family (Werfel et al., 2003; Widrow & Lehr, 1990) rather than weight-perturbation ES. The key structural innovation is the virtual population concept: because perturbations happen at activation space per-token, a single forward pass evaluates a population proportional to the sequence length, rather than requiring separate forward passes per population member. This is what buys the 10³–10⁴× efficiency gain over weight-space ES. The method introduces two biases: different credit assignment rules for different layer types (attention vs. linear), and interference avoidance between perturbed modules. These are modest inductive biases compared to backprop's requirement of end-to-end differentiability. The integrity picture is mixed but appropriately honest. The authors compare against backprop (the real SOTA) and EGGROLL (the strongest ES baseline). They acknowledge the compute gap explicitly and do not oversell. However, the EGGROLL comparison relies on extrapolations rather than head-to-head runs at identical scale, which introduces uncertainty. The paper tests up to 243M parameters and 1B tokens — serious but not frontier scale. No pre-registration. Code is referenced. The key risk is that the favorable scaling properties observed at 243M may not hold at 1B+ parameters, and we have no data points there yet. The paper is deliberately honest about what it leaves on the table. It does not train non-differentiable architectures — the exact class of models that would uniquely benefit from zeroth-order training. It does not attempt compute efficiency. These are the two experiments that would transform Dust from 'interesting proof of concept' to 'paradigm shift,' and the authors explicitly flag both as future work. The most likely read is (a) — they're establishing the foundation and ran out of scope — but (c) is also plausible: training a non-differentiable architecture with Dust would be the obvious next paper.