Imagine you install a thermostat that can shut off rooms individually when they're warm enough — a smarter system than the old "always heat two rooms" approach. But after a few weeks, you notice the boiler has learned to run cooler overall, so the thermostat's shutoff threshold is almost never triggered. The rooms stay heated. The smart thermostat is still there, but the system has routed around it. That's what happens when you swap softmax for sparse probability maps in mixture-of-experts language models. The committed claim: replacing softmax with sparsity-inducing maps (sparsemax, entmax, normmax) in top-K MoE routers does NOT automatically produce adaptive expert participation — because the router's learned score distribution co-adapts to the map, and the map's theoretical capacity to zero out experts is largely neutralized by training dynamics. This is a negative result dressed as a mechanistic explanation, and it's more useful than most positive results in this space. The authors train matched 300M and 1B parameter top-2 MoE language models across four maps and find sharply divergent behavior. At 1B scale, entmax discards 30% less probability mass than softmax while almost never dropping a selected expert. Sparsemax retains the most mass. Normmax routes 21% of tokens to a single expert — the only map that genuinely achieves adaptive participation, and it does so by learning a different score gap distribution, not because of an intrinsic map property. The key insight is geometric: each map drops a second expert only when the gap between the two largest router scores crosses a fixed, map-specific threshold, and the router learns score distributions that stay mostly below or above that threshold. The ladder position is honest and a bit deflating. None of the sparse maps improve validation loss over softmax. The win is elsewhere: robustness to inference-time expert count changes. Sparsemax trained with K=2 loses only 0.02 nats when run with K=8, versus 0.58 nats for softmax. That's a 29× gap in inference flexibility. This matters for deployment — you want to be able to dial expert count up or down without retraining — but it's an operational advantage, not a modeling one. The architecture story is clean. This is standard Transformer-based MoE with token-choice top-K routing, swapping only the probability map applied to router logits. The novelty is entirely in the analysis of what the map change induces in the learned score distribution. The paper treats the map-score interaction as the unit of analysis rather than the map alone — a framing the field needs. The theoretical contribution (Propositions 4.1–4.3 characterizing each map's threshold geometry) is the load-bearing structure. Integrity is solid for this kind of work. The authors compare four maps at two scales with matched hyperparameters, report perplexity and auxiliary metrics transparently, and show that their theoretical thresholds predict empirical behavior. The limitation is that this is entirely self-contained simulation — no external benchmark, no downstream task evaluation, no independent replication. The scale (300M and 1B) is useful for mechanistic study but leaves open whether these dynamics hold at frontier scale. The successor question is obvious: what happens at 7B+ or with expert-choice routing instead of token-choice? The authors study top-K token-choice exclusively. Expert-choice routing (where experts pick tokens, not vice versa) has different load-balancing dynamics that could interact with sparse maps in entirely different ways. My read: this is (a) compute-constrained and (c) saved for the next paper. The 1B scale is already nontrivial for an academic group, and expert-choice routing would require a second full set of experiments.