Imagine you're adjusting a recipe that has twenty sequential steps — marinate, sear, deglaze, braise, and so on. If you want to tweak the final flavor (the reward), the textbook approach is to trace backward through every step, asking: how did changing the salt at step 3 affect the fond at step 7, which affected the reduction at step 15, which changed the flavor at step 20? That chain of dependencies is the vector-Jacobian product in adjoint matching. Now imagine you discover that, in practice, each step's influence on the next is almost entirely one-to-one — the salt mainly affects saltiness, the acid mainly affects acidity, each on its own channel. If that's true, you don't need the full chain; you can just scale the final flavor gradient by how many steps remain. That is exactly what this paper discovers and exploits. The committed claim: the batch-averaged velocity Jacobian of pretrained flow policies concentrates on its diagonal, which permits replacing the full vector-Jacobian product at every flow step with a single scalar multiplication by the flow time. This eliminates the dominant cost of adjoint matching — the per-step backpropagation through the policy — while preserving the quality signal from the learned critic. The result is a method called SQAM (Q-learning with Scalar Adjoint Matching) that is dramatically cheaper to run and, on the hardest tasks, dramatically better. The ladder is clear and honestly reported. The authors benchmark against a strong field on OGBench, the current community standard for offline goal-conditioned RL. SQAM's gains are concentrated on the four hardest domains — not cherry-picked easy wins — where it exceeds the strongest per-domain baseline by 18 to 35 percentage points in success rate. This is a large margin in a benchmark where many methods plateau. Importantly, on easier domains, the improvements are modest, and the authors do not hide this. They also compare against full (non-scalar) adjoint matching, showing the scalar approximation loses little accuracy while cutting computational cost substantially. Architecturally, SQAM lives in the flow-matching family — continuous normalizing flows that generate actions by integrating a learned velocity field over time steps. The key insight is structural: if the Jacobian of that velocity field is near-diagonal in expectation, you can collapse the ODE adjoint into a scalar. This is not a universal property of flow models; it is an empirical observation about pretrained flow policies specifically. The method also introduces a value penalty at policy-generated actions, a secondary but important contribution that stabilizes training under the scalar approximation. The compute savings are real: no per-step vector-Jacobian products means the cost no longer scales with the number of flow steps. On integrity, the validation is solid but not ironclad. OGBench is a recognized community benchmark — not a bespoke suite — and the baselines include current competitive methods (IDQL, HIQL, QRL, among others). The robot experiments on a real bimanual platform add a crucial layer beyond simulation, though with only three tasks and no independent replication yet. The diagonal-concentration claim is supported empirically with spectral analysis, but the theoretical conditions under which it holds (or breaks) are not fully characterized. This is an honest gap rather than a hidden one. The milestone to track: SQAM works on a vision-language-action (VLA) policy on a real robot, but the scale is still modest — three tasks, one robot. The next meaningful threshold is whether this approach holds at the scale of foundation-model-class VLAs (hundreds of tasks, multiple embodiments) without the diagonal assumption degrading. If it does, scalar adjoint matching becomes a default fine-tuning primitive for robotics foundation models. The obvious experiment not run: fine-tuning a truly large VLA (e.g., RT-2 scale, billions of parameters) on a diverse multi-task suite where the diagonal assumption might break. The honest read is (a) — compute and robot access. Training a billion-parameter VLA with RL in the loop is expensive, and the authors likely had access to one robot platform. The three-task demonstration is a proof of concept, not a scaling study, and they know it.