Imagine you have a GPS that gives you perfect directions — if you let it recalculate every 50 meters. Now you need it to give you one instruction at the start of the highway on-ramp and still get you to the right exit. That's the problem RFPO solves for robot control: flow-based policies generate smooth, expressive action trajectories by integrating an ODE over many steps, but on a real robot you can't afford 64 integration steps per action. You need one step, and the policy needs to still work. The committed claim: RFPO is the first framework to explicitly close the "few-step discretization gap" in flow-based robot policies, preserving full-step control quality under single-step execution across four distinct robot morphologies. This isn't a speed-accuracy tradeoff paper — the authors argue the tradeoff itself is largely eliminable with the right training recipe. The architecture sits in the flow-matching / rectified-flow family — think of it as a cousin to diffusion policies but with ODE-based transport instead of stochastic denoising. The key trick is "reward-aware online Reflow": during on-policy training, a student policy learns to straighten the transport paths so that a single Euler step approximates what 64 steps would produce. A frozen Gaussian PPO controller provides complementary supervision at full and intermediate integration budgets, but the deployed policy is just the student running one step. The training is more expensive; the deployment is dramatically cheaper. The numbers are striking and consistent. On Unitree Go2, one-step execution retains 98.5% of 64-step reward while cutting mean inference latency from 4.39 ms to 0.08 ms — a 54.9× speedup. Across Go2, Spot, H1, and G1 robots, the one-step return stays within 2.4% of the 64-step value under both zero and random initialization. These aren't cherry-picked best cases; the consistency across four morphologically distinct platforms is the real signal. On the integrity side, validation spans simulation across four robot platforms plus real-robot locomotion experiments on Go2. The baselines include standard flow policies at various step counts and the PPO teacher itself, which is reasonable for establishing the discretization gap but doesn't pit RFPO against the broader field of fast-inference policy architectures (e.g., standard MLPs, one-shot diffusion distillation methods). The real-robot experiments add credibility but are locomotion-only — no manipulation, no contact-rich tasks. Code is released. The milestone that matters is generalization beyond locomotion. Locomotion is the friendliest domain for this kind of distillation because the action space is relatively smooth and cyclic. The real test is whether RFPO's one-step execution holds up in contact-rich manipulation, where the ODE landscape is far less forgiving. The authors tested four legged robots but zero arms — that's the gap. The obvious next experiment: dexterous manipulation on a multi-fingered hand. The flow policy's expressiveness matters most when the action distribution is multimodal (reach around vs. reach over an obstacle), and that's exactly where one-step approximation is hardest. The honest read: this is being saved for the next paper. The framework is set up for it, the four-robot breadth suggests the authors are methodical, and manipulation is the natural sequel that would elevate RFPO from a locomotion trick to a general-purpose tool.