You know how a good coach watches game tape? They don't redesign the player — they diagnose what went wrong in the last play and give a specific correction: 'You're gripping too hard on the backhand.' PEARS does exactly this for robots. After a failed manipulation episode, a vision-language model looks at what happened (the visual outcome, the force history from the tactile sensor) and reasons about the physics: was the contact force too high? Was timing off? It then updates explicit force bounds — not the entire policy, just the guardrails. The committed claim: a pretrained robotic manipulation policy can be adapted online to out-of-distribution conditions using dramatically fewer real-world trials by combining physics-guided failure diagnosis (via VLM) with latent-space steering of a frozen diffusion policy. The paper reports 12.4–37.4 percentage point improvements over per-task baselines in simulation and up to 53.2% fewer episodes needed to hit a target success rate. The architecture is a hybrid of two distinct loops operating at different frequencies. The physics-guided force reasoning (PFR) module is an episodic, low-frequency loop: after each trial, a VLM diagnoses failure from visual and tactile data and adjusts contact-force bounds enforced by a high-frequency hybrid force-position controller during the next episode. Complementing this, a tactile-conditioned diffusion steering RL module operates at decision time — it perturbs the latent noise of a frozen flow-matching policy to correct free-space trajectory and contact timing errors without updating any weights in the base model. This is the core architectural bet: separating contact-force correction (physics priors, explicit bounds) from trajectory correction (latent-space RL steering) and letting each module handle what it's best at. The ladder here is encouraging but demands careful reading. In simulation, PEARS beats the strongest per-task baselines by 12.4–37.4 percentage points on success rate and cuts required interactions by up to 53.2%. Real-world results — 95% on whiteboard erasing, 90% on pipette aspiration — are strong but come from just two tasks. The baselines compared in simulation include ablations (PFR-only, steering-only) and prior online adaptation methods, which is honest design. However, we don't see a direct comparison to the very latest diffusion-policy adaptation methods from 2025, and the real-world experiments are limited in task diversity. Integrity is solid for a robotics paper at this stage. The simulation includes multiple tasks with ablations that isolate each module's contribution — a critical design choice that prevents the 'was it the VLM or the steering?' ambiguity. Real-world validation on physical hardware is present, which immediately elevates this above sim-only work. But the real experiments cover only two tasks, neither involving high-precision multi-object manipulation. No pre-registration, no independent replication — standard for the field but worth flagging. The milestone question is about sample efficiency thresholds. Right now we're at tens of episodes for adaptation. The unlock everyone is watching for is single-digit episode adaptation on dexterous multi-object tasks — think assembling a connector or handling deformable materials — where each failed trial has real cost (broken parts, wasted reagents). PEARS gets the episode count down by ~53%, but the absolute numbers aren't stated for real-world tasks, making it hard to pin down exactly where on the curve we sit. The obvious experiment not run: multi-object or deformable-object manipulation in the real world, where force reasoning gets vastly more complex (multiple simultaneous contacts, nonlinear material response). My read is (a) and (c) — the hardware and experiment setup for deformable manipulation is genuinely expensive and slow, and this is almost certainly the next paper's territory. The two real-world tasks chosen (erasing, pipette aspiration) are both single-contact-surface tasks, which is where PFR's force-bound reasoning is cleanest.