Imagine you're learning to cook from a recipe book, but every time a dish fails you can rewind time, watch a slow-motion replay from every angle, read the chef's margin notes about what the ingredients were actually doing, and then rewrite the recipe before trying again. That's RPG — Reconstruct, Practice, Go Real — a framework where a robot manipulation system improves itself by building practice arenas from offline demonstrations, diagnosing its own failures using privileged simulator information, and iteratively rewriting the symbolic skills and system prompts that drive its behavior. No neural network retraining. No gradient descent. Just code-level self-debugging in a loop. The committed claim: a robot execution system can autonomously improve from 28.6% to 95.0% task success across 22 held-out manipulation tasks over 15 practice rounds, without updating any model weights, and then transfer frozen to physical hardware at 100% (30/30 trials on 3 tasks). This beats ASPIRE at 75.5% and CaP-Agent0 powered by GPT-6 Astra Pro at 60.0%. The core insight is that LLM-based robot systems fail not because the foundation models are bad, but because the symbolic skill libraries, prompts, and coordination logic surrounding them are undertrained — and you can fix those with structured self-practice rather than expensive data collection or fine-tuning. The architecture is distinctive: RPG operates entirely in the outer loop around a frozen multimodal LLM. It takes an offline dataset of demonstrations, reconstructs simulation environments that match them, then generates practice tasks related to the target capabilities. During practice, the system has access to three diagnostic channels — execution feedback (did the gripper close?), privileged simulator state (where is the object actually?), and the original demonstration videos. When a task fails, RPG uses these channels to localize the failure to a specific skill, perception call, or prompt instruction. It then proposes candidate fixes: new reusable symbolic skills, refinements to existing skills, or revisions to the system prompt. Crucially, cross-task evaluation gates every change — individual candidates and merged revisions are tested across tasks before being retained, preventing regression. The ladder comparison is strong and honest. RPG doesn't just beat weak baselines — it names ASPIRE (75.5%), CaP-Agent0 with GPT-6 Astra Pro (60.0%), and initial CaP-Agent0 (28.6%) as comparisons across the same 22-task benchmark. The 95.0% number after 15 rounds represents a 3.3× improvement over the starting point. The physical transfer result — 30/30 on three tasks after a calibration procedure — is small-scale but clean. The gap between RPG and the GPT-6-powered baseline is particularly telling: a better foundation model alone (GPT-6 Astra Pro) gets you to 60%, but structured self-practice on top of a weaker system gets you to 95%. The implication is that the bottleneck isn't model capability but skill-library quality. Integrity is reasonable for a robotics paper but has the usual caveats. The 22-task benchmark is held-out initializations, which is better than training-set evaluation, but the tasks themselves are chosen by the authors. Physical trials are limited to 3 tasks × 10 trials each — enough to demonstrate transfer but not enough to claim generalization to arbitrary real-world conditions. The calibration and hardware-adaptation procedure is described as 'common' but its scope matters: if it's substantial, the 'frozen system' claim is weaker than it sounds. No independent replication exists yet, and the code/environment availability isn't explicitly stated in the abstract. The milestone framing here is about closing the sim-to-real gap for LLM-driven manipulation. Today: 95% in sim across 22 tasks, 100% on 3 physical tasks with 10 trials each. The next meaningful number is something like 90%+ success on 50+ diverse physical tasks without per-task calibration — that's the threshold where you'd trust this in a real deployment setting. The gap is probably 2-4 years, gated less by the RPG framework itself and more by simulation fidelity, task diversity, and the robustness of the underlying perception and control primitives. The obvious experiment not run: scaling to tasks that require long-horizon planning with more than 5-10 steps, tasks with deformable objects, or tasks requiring force-sensitive manipulation. The honest read is (a) — simulation fidelity for these domains isn't there yet, and RPG's reconstruct step depends on being able to build faithful practice environments. The authors likely know this and are saving the deformable-object and multi-step generalization story for subsequent work once simulation tools catch up.