Imagine you're debugging a complex codebase you've never seen before. You could read every file top to bottom — or you could form hypotheses about where the bug lives, write a test that distinguishes between your two best guesses, run it, and let the result tell you which branch to prune. EmbodiedRSI applies exactly this logic to robot learning: instead of grinding through demonstrations or random rollouts, the system maintains competing hypotheses about how to solve a task, then selects the single physical experiment most likely to resolve the disagreement. The committed claim: a robot harness that autonomously chooses where to explore, runs physical trials, and co-evolves its own code and skills — reaching 77.0% success on RoboCasa365 overall and 71.3% on Composite-Unseen tasks, compared with 40.1% for the best existing baseline. That is not an incremental improvement; it is a near-doubling of success rate on tasks the system has never encountered. The architecture is a Fast-Slow Dual-System, loosely inspired by Kahneman's System 1 / System 2 distinction. The Fast System executes skills and code rapidly. The Slow System maintains a Hypothesis Graph — a structured representation of competing code and skill variants — and builds Hierarchical Memory from accumulated experience. Value-of-Information Experiment Selection is the key mechanism: rather than exploring randomly or exhaustively, the system computes which physical trial would most reduce uncertainty across its hypothesis set, then runs only that trial. This is classic Bayesian experimental design applied to robot self-improvement, but the novelty is making it work end-to-end in an agentic code-generation loop. The ladder position is strong on the benchmarks shown. On RoboCasa365, the 77.0% overall and 71.3% Composite-Unseen numbers substantially exceed the 40.1% best baseline. On LIBERO-Pro, 86.8% overall success is similarly dominant. The real-world transfer result — 71.3% across multiple challenging tasks with zero-shot transfer — is the load-bearing number. However, the abstract does not name the specific baselines by method, which makes independent assessment harder. We know the gap is large; we don't know exactly who lost. Integrity is mixed. The system is tested on two established benchmarks (RoboCasa365 and LIBERO-Pro), which is good practice. Real-world transfer adds a physical validation layer beyond pure simulation. But we don't see pre-registration, and the abstract doesn't clarify whether benchmark selection was decided before or after results. Code availability is not mentioned. The 365-task scale of RoboCasa365 is meaningfully large for this domain, but the specific conditions of 'Composite-Unseen' — what exactly is unseen and how it was partitioned — matters enormously for interpreting the 71.3% number. The milestone question is about scaling: can this approach maintain its advantage as task complexity grows beyond kitchen manipulation toward long-horizon, multi-room, multi-object tasks? The gap between 77% and human-level reliability (~95%+) is where the real engineering happens. Closing the last 20% on unseen compositional tasks is historically where robot learning systems stall. The obvious experiment not run: adversarial task generation. The system selects its own experiments to resolve its own hypotheses, but what happens when the environment is actively hostile to those hypotheses — novel objects, deformable materials, dynamic obstacles that weren't in any training distribution? The authors likely stopped at kitchen-scale because that's where the benchmarks live, and extending to truly open-ended environments would require a fundamentally different evaluation infrastructure. This is probably (a) — compute and infrastructure limits — but it's the experiment that would tell us whether Value-of-Information selection generalizes or overfits to the hypothesis space the system can represent.