Imagine you're training for a marathon by running laps on a track, but your coach randomly assigns you either a 100-meter jog or an ultramarathon each session. You'd crush the easy laps without learning anything and fail the impossible ones without useful feedback. What you actually need is someone watching your fitness edge — the distance where you succeed about half the time — and scheduling runs right there. That's what Success Guided Sampling (SGS) does for reinforcement learning in simulation. The committed claim: when you scale RL to a million parallel environments with diverse simulator resets, uniform sampling over task configurations wastes most of your batch on tasks the policy already solved or cannot yet attempt. SGS fixes this by concentrating training on the policy's capability frontier, and the fix is large enough to unlock tasks — multi-terrain quadruped locomotion, contact-rich peg-in-hole assembly — that prior uniform-sampling approaches simply fail to solve, even at the same massive scale. The method sits inside the "massively parallel sim-to-real RL" pipeline popularized by work like Rudin et al. (Legged Gym) and Tao Chen et al.'s recent diverse-reset paradigm for manipulation. Those pipelines run on GPU-accelerated simulators (IsaacGym / IsaacLab), and their core bottleneck isn't compute per se — it's that as you scale environments, the fraction of useful learning signal in each batch drops because most environments are set to configurations the policy already handles or can't touch. SGS partitions task configurations into bins, tracks per-bin success rate, and upweights bins near the ~50% success frontier using a simple peaked sampling distribution. It's lightweight — a histogram update per batch, no learned curriculum model. The ladder matters here. The direct comparison is against the Tao Chen et al. (2024) uniform-reset baseline at identical environment counts. On quadruped locomotion across mixed terrains, uniform sampling flatlines at partial terrain coverage while SGS solves the full suite at 2^20 environments. On peg insertion, uniform sampling fails entirely at tight tolerances (0.5mm) where SGS succeeds. The authors also compare against domain randomization and manual curriculum baselines; SGS matches or beats all of them without per-task tuning. Critically, SGS's advantage grows with scale — at small environment counts the gap is modest, but at 2^18–2^20 the uniform approach hits a wall while SGS keeps climbing. Integrity is mixed but honest. All results are in simulation (IsaacLab), graded by the same team that built the method — the classic "grading your own homework" setup. However, the authors do transfer learned manipulation policies to real hardware via RGB-based distillation and demonstrate zero-shot success on physical assembly tasks, which is the strongest form of validation available for sim-to-real work. No pre-registration, no independent replication yet, but the method is simple enough that reproduction is tractable. Code and project website are provided. The milestone question is concrete: SGS enables 0.5mm-tolerance peg insertion in sim with zero-shot real transfer. The next meaningful number is sub-0.1mm tolerances or multi-step assembly sequences (e.g., NIST assembly benchmarks with 5+ parts), which would move this from research demo to industrial relevance. The gap is probably 2-3 years and depends as much on sim fidelity and contact modeling as on the curriculum method itself. The obvious experiment not run: scaling SGS to deformable-object manipulation or tasks with truly discontinuous reward landscapes (e.g., tool use, cloth folding) where the success-rate frontier may not be smooth. My read is (a) — these tasks require simulators the team didn't have or didn't trust, not that the method was tried and failed. The second missing piece is a head-to-head against learned automatic curriculum methods (like PLR or PAIRED) at this scale; the authors compare against simpler baselines but not the strongest curriculum-learning competitors.