Imagine you build a Waze for school admissions. You tell every family which schools are great and nearby and likely to accept their kid. Brilliant — except now everyone turns down the same street, and the shortcut becomes gridlock. The "best route" was only best when nobody else knew about it. This paper formalizes that exact failure mode and builds a recommender that accounts for the traffic it creates. The committed claim: naive recommendations in capacity-constrained matching markets cause congestion that disproportionately harms applicants with the fewest alternatives, and a congestion-aware bilevel optimize-and-simulate approach can safely improve match outcomes in equilibrium. The authors don't just prove this theoretically — they deployed it in the 2025-26 NYC high school admissions cycle with a randomized controlled trial covering real eighth graders making real decisions about their futures. The architecture is a bilevel optimization problem. The outer level allocates personalized recommendation lists across applicants. The inner level simulates the deferred-acceptance matching mechanism that NYC actually uses, modeling how applicants respond to recommendations. The key structural choice is treating recommendations as a market-shaping intervention rather than a pure information provision — the recommender must anticipate its own equilibrium effects. This places the work squarely in the mechanism design tradition, borrowing from both the matching markets literature (Gale-Shapley, Roth) and the algorithmic fairness literature on disparate impact. The RCT results are promising but appropriately sized for a first deployment. Treatment applicants ranked a recommended program at a 16.4% rate versus 10.5% for controls — a 57% relative increase (p=0.011). On the harder outcome of actually matching to a recommended program, treatment hit 5.6% versus 3.3% for controls — a 71% relative increase, though with p=0.071, which sits just outside conventional significance. Critically, zero treatment applicants were rejected from a recommended program, which validates the congestion-aware design's core safety guarantee. The integrity picture is strong for a first-deployment paper. This is a real RCT embedded in an actual high-stakes admissions process — not a simulation grading its own homework. The comparison is treatment-versus-control with a clear pre-specified primary outcome. The weakness is the p=0.071 on the match outcome — the authors are transparent about this, but it means the headline "71% relative increase in matches" is suggestive rather than confirmed. The sample size was constrained by the deployment context, not by choice. The ladder question is interesting because the baseline isn't another algorithm — it's the status quo of no personalized recommendations at all. The paper's real comparison is against naive (non-congestion-aware) recommendations, which they show theoretically cause sharp acceptance-rate decreases. There's no head-to-head against another deployed congestion-aware recommender because, as far as the literature is concerned, this is the first one deployed at this scale in a real matching market. The successor experiment that wasn't run: a multi-year longitudinal study tracking whether recommended matches lead to better educational outcomes (graduation rates, satisfaction, college enrollment). The authors almost certainly know this is the question that matters most, but it requires years of follow-up data they don't have yet. This is a classic case of (a) — the data simply doesn't exist yet, not a strategic omission. The other obvious missing experiment is scaling beyond a single admissions cycle to test whether the congestion-aware properties hold as the recommender reshapes applicant behavior over repeated years.