Imagine you've spent years learning to play piano — scales, chords, arpeggios, the full repertoire of finger mechanics. Now someone hands you sheet music for a piece you've never seen. You don't need to re-learn the piano. You need someone to tell you which notes to play, in what order, and with what dynamics. If the sheet music is wrong, you sound terrible despite having perfect technique. InterEvolve is a system for writing better sheet music for a humanoid robot whose fingers already know the piano. The committed claim: a pretrained humanoid loco-manipulation controller already contains far more motor competence than any single hand-designed reward can unlock, and that hidden competence becomes accessible through an LLM-guided evolutionary search over reward programs at test time — no retraining required. The system composes a forward-backward (FB) behavioral foundation model (the pianist's hands) with structured reward programs (the sheet music) whose code structure is revised by an LLM agent while a numerical optimizer tunes the constants. Each candidate program is verified across parallel simulation scenarios before being accepted into a growing skill library. The architecture has two load-bearing pieces. First, an object-aware forward-backward model that adds object residuals on top of a frozen body-motion prior. This is the key representational choice: the FB model encodes a continuous space of behaviors reachable by reward conditioning, so the reward program becomes the interface between task specification and motor execution. Second, the reward programs themselves are staged rewards with completion conditions and tunable constants — code, not vectors. The LLM revises program structure (which stages, which conditions, which terms), while CMA-ES tunes the numerical constants. This separation of structure search from parameter search is the core engineering insight. Where does this sit on the ladder? The paper's main comparison is against human-designed reward functions paired with the same FB controller. The headline result is that InterEvolve's evolved rewards substantially outperform expert-written ones, demonstrating that manual reward engineering leaves significant motor competence on the table. They also compare against prior reward-search methods (Eureka, DrEureka) and show gains. However, the baselines are all within the reward-shaping paradigm — there is no comparison against end-to-end RL retrained per-task or against model-predictive control approaches. The comparison is honest within its frame but the frame is narrow. Integrity checks reveal a mixed picture. Validation is entirely in the authors' own simulation (Isaac Gym / Isaac Lab) plus a single-robot physical deployment on a Unitree G1. The physical demos are compelling but anecdotal — no systematic real-world benchmark. The simulation tasks are diverse (object manipulation, locomotion, multi-stage compositions) but self-selected. There is no pre-registration and no community benchmark for humanoid loco-manipulation at this level, which is partly a field maturity issue. Code and project page are provided. The milestone to watch is clear: test-time adaptation that works reliably in unstructured real environments, not just sim-to-real transfer of individually evolved skills. Today the system evolves programs in simulation and then deploys; the gap is closing that loop so the robot adapts from real-world execution feedback. The paper demonstrates sim-only evolution with real deployment — the next concrete step is real-world feedback closing the evolution loop, likely requiring 10-100x faster sim-to-real iteration cycles. The obvious experiment not run: scaling to genuinely novel object categories and contact geometries not represented in the FB model's training distribution. The FB model was trained on a specific set of scenarios; the paper claims test-time generalization but every demonstrated task plausibly falls within the training distribution's convex hull. The authors likely know this boundary exists — pushing past it would require either a much larger pretraining distribution or architectural changes to the FB model itself. Best read: they're saving the distribution-boundary investigation for the next paper, where it becomes the central question.