You know those assembly-line quality inspectors who watch parts roll past and pull defective ones before they reach the end of the line? They don't wait for the finished product to fail — they watch the shape of the process. If the third weld looks wrong, they yank it. OnTrack does this for LLM agents: instead of waiting for an agent to finish a 40-step run and then grading the result, it compares each step's trajectory against known-good runs and flags divergence in about a millisecond. The committed claim: a streaming, structure-aware optimal transport mechanism can monitor LLM agent trajectories in real time, detecting failures early enough to abort before compute is wasted — and it does this faster and more accurately than content-similarity baselines. This is not a new agent architecture or a new reasoning technique. It is plumbing. Specifically, it is the plumbing that lets you deploy agents without crossing your fingers. The method works across three regimes of decreasing data access, which is where the engineering honesty lives. With full reference access — historical successful runs plus tool schemas — OnTrack can detect plan violations by comparing an agent's dependency graph against reference graphs using structure-aware optimal transport. With only tool schemas, it can still catch structural anomalies. With nothing but step logs, it degrades gracefully to loop detection, stall identification, and repeated-tool-call flagging. The paper is explicit that monitoring capability drops as access drops. That honesty matters. The evaluation is on SWE-bench trajectories, which is a real community benchmark for software engineering agents. Using only the first 8 steps of a run, OnTrack ranks failing trajectories below succeeding ones with +0.057 AUROC improvement over content-similarity approaches. That delta is modest. The more operationally meaningful number: an abort policy built on OnTrack saves roughly 18% of compute that would otherwise be burned on failing runs, with 83% precision — 5 out of 6 aborted runs were genuinely heading to failure. The architectural family here is optimal transport (OT), specifically applied in a streaming fashion to graph-structured agent traces. This is not gradient-based learning or fine-tuning. It is a comparison mechanism: you have a reference distribution of successful trajectory structures and you measure how far the current run's structure has diverged. The key structural choice is making OT structure-aware — not just comparing step content but the dependency graph between steps and tools. The compute property it leans on is cheapness: ~1ms per step means this sits in the critical path without meaningful latency cost, unlike a safeguard LLM that adds inference time at every step. The baseline comparison is the weakest point. Content similarity is the named competitor, and +0.057 AUROC is a real but thin margin. The paper does not compare against other real-time monitoring approaches — though the authors argue that real-time structural monitoring of this kind barely exists as a category, which is partially true. The SWE-bench evaluation is a strong benchmark choice, but the sample of abort decisions (6 total aborts, 5 correct) is small enough that the 83% precision number carries wide confidence intervals. The field fight this participates in is whether agent safety should be handled by another LLM (expensive, slow, recursive trust problem) or by lightweight structural monitoring (cheap, fast, but limited understanding). OnTrack is a strong argument for the structural-monitoring side, at least for the failure modes it can catch. The obvious next experiment is scaling this to longer trajectories, multi-agent systems, and domains beyond software engineering — and the honest read is that the authors are saving those for follow-on work, not that they tried and failed.