Imagine you're training a new chef by having them watch cooking videos. Some videos use a professional kitchen with induction burners and convection ovens; others use a camping stove and a pocket knife. If you mix those videos together without labeling which kitchen the chef will actually work in, the chef learns contradictory muscle memory — reaching for equipment that isn't there, timing recipes for heat sources that don't exist. CoTrace is the realization that the kitchen matters as much as the recipe, and you need to tag every training video with which kitchen produced it. The committed claim: trajectory data for training terminal agents (think: AI that runs shell commands to solve coding tasks) is not fungible across runtime harnesses. A harness is the scaffolding that formats prompts, binds tools, and handles error recovery — the invisible exoskeleton around the language model. CoTrace introduces provenance-matched data recipes that route training trajectories only to models running under the harness that generated them, and this discipline produces large gains at substantially lower compute than naively pooling all trajectories together. The alternating co-evolution framework is clean engineering: harness search and policy training are decoupled into separate promotion decisions. Recurring execution failures during rollouts feed back into harness synthesis — the scaffolding literally learns from the model's mistakes. Meanwhile, policy training via supervised fine-tuning (SFT) only sees verified rollouts matched to the current runtime, and the RL variant uses fresh online interactions rather than stale replay. On the Tmax promotion split, this takes Qwen3.5-9B from 78 solved tasks (baseline) to 88 under SFT and 90 with online RL. The key finding isn't just the top-line number — it's that a compact, harness-matched corpus beats a much larger corpus pooled across sibling harnesses. Data quality via provenance tracking beats data volume. The ladder comparison is where this gets interesting. The paper evaluates on Terminal-Bench 2.1 and SWE-bench Lite, which are community benchmarks with real adoption. The out-of-distribution transfer results are the sharpest contribution: when you evaluate a model trained under one harness using a different harness at test time, performance collapses. This isn't a subtle effect — the paper calls them 'procedural execution breakdowns.' The implication is that the entire field of terminal-agent benchmarking has a confound: you can't meaningfully compare two agents unless you control for harness compatibility. On integrity, the paper is solid but not airtight. Tmax is an internal split, not a pre-registered community benchmark. Terminal-Bench 2.1 and SWE-bench Lite are public and well-adopted, which helps. The paper is dense — 32 pages, 7 figures, 17 tables — suggesting the authors aren't hiding unfavorable ablations. But there's no independent replication, and the harness search space is defined by the authors, leaving open the question of how well this generalizes to harnesses designed by other teams. Code availability is not explicitly confirmed in the abstract. The milestone question is about scale and generality. Today we're at 88-90 solved tasks on Tmax with a 9B model. The next meaningful number is whether this recipe transfers to 70B+ models and whether harness-aware training can push SWE-bench Lite scores past the current frontier without ballooning compute. The broader milestone is whether harness provenance becomes a standard field in training data pipelines — if it does, this paper will be cited as the reason. The obvious experiment not run: training with deliberately adversarial harness mismatches to quantify the degradation curve precisely, and testing whether a model trained on provenance-matched data from harness A can be cheaply adapted to harness B via a small matched fine-tuning step. The authors likely ran out of compute budget — the co-evolution loop is already expensive — but this transfer-learning question is the natural next step and would dramatically increase practical value.