Imagine you're a sculptor working with wet clay on a spinning wheel. You have two sources of guidance: a photograph of the target shape (the demonstration data), and a friend watching from the side shouting 'push left, no — too much, back right' (the critic). Most prior approaches either let you study the photo endlessly before touching the clay (on-policy, slow) or let your friend shout directions but risk the clay flying off the wheel when the advice is wrong (off-policy, unstable). QF3's trick: only listen to your friend's corrections on the dimensions where your hands are already close to where they should be. Ignore the rest. That filtered trust is what makes it work. The committed claim: QF3 is the first off-policy flow RL method to train humanoid locomotion policies from scratch and transfer them zero-shot to real hardware. Flow policies — continuous-time generative models that learn to map noise to actions — have become the default policy class for learning from demonstrations, but reinforcing them through interaction has been brittle. On-policy methods like FPO++ work but are sample-hungry and slow. QF3 breaks this by going off-policy: it trains a flow policy using flow matching loss plus backpropagated critic gradients through a single-step prediction of the flow's output, filtered so only action dimensions near the replay buffer action receive the gradient. This dimensional filtering is the architectural novelty — it prevents the critic from pushing actions into regions where neither the critic nor the one-step flow prediction is reliable. On the ladder: QF3 is benchmarked against FPO++ (the recent on-policy flow RL method), standard DPPO, and diffusion-policy baselines. The headline number is a 10× wall-clock speedup over FPO++ for humanoid locomotion and motion-tracking tasks. It also fine-tunes pretrained flow-based manipulation policies on ABC-Sim and Robomimic benchmarks. The paper claims competitive or superior performance, though the exact reward curves and final-performance deltas versus classical RL baselines (SAC, TD3) on the same tasks are what the community will scrutinize. The zero-shot sim-to-real transfer on a real humanoid is a strong existence proof, not yet a controlled comparison against tuned classical pipelines. Architecturally, QF3 sits in the actor-critic family, specifically the deterministic-policy-gradient lineage (DDPG → TD3 → DPPO), but replaces the deterministic actor with a flow-matching generative model. The key structural choice is the one-step prediction shortcut: rather than rolling out the full ODE of the flow during training, QF3 backpropagates the critic gradient through a single Euler step of the flow's velocity field. This trades some expressiveness for computational tractability and gradient stability. The dimensional filtering — masking critic gradients to action dimensions within a threshold of the replay action — is a regularization trick that keeps the optimization in-distribution. The whole thing runs on a high-throughput off-policy training recipe, meaning large replay buffers and parallel simulation. Integrity is mixed. The paper demonstrates on standard sim benchmarks (IsaacGym humanoid, ABC-Sim, Robomimic) and shows a real-hardware transfer video, which is meaningful but not independently replicated. Baselines include FPO++ and DPPO but not the full classical off-policy zoo (SAC, TD3, TQC) on identical tasks, which leaves a gap. The 10× speedup is wall-clock, which is hardware- and implementation-dependent — a fair metric for practitioners but not a clean algorithmic comparison. No pre-registration, no independent replication yet. Code availability via the project page will be the key integrity signal. The milestone to watch: QF3 demonstrates locomotion and motion-tracking on a single humanoid morphology. The next concrete threshold is multi-task, multi-morphology flow RL — training a single flow policy that generalizes across humanoid, quadruped, and dexterous manipulation tasks without retraining. That's probably 1-2 years out if the architectural pattern holds. The broader milestone for the field is real-world dexterous manipulation from scratch via off-policy flow RL, which would require handling contact-rich dynamics and much longer horizons than locomotion. The obvious experiment not run: training QF3 head-to-head against SAC/TD3 on the exact same humanoid tasks with identical compute budgets. The honest read is (c) — the authors are positioning flow policies as the future policy class and want to show flow-vs-flow comparisons (QF3 vs FPO++) rather than flow-vs-classical, likely saving the classical showdown for a follow-up or leaving it to the community. This is a reasonable strategic choice but it means the paper doesn't answer the question most practitioners will ask: 'Why should I use a flow policy instead of SAC?'