Imagine you're cooking a complex multi-course dinner from a recipe book, but you can only keep a few pages open at a time. Early on, you need the stock recipe — but once the stock is made, keeping that page open wastes precious counter space. A skilled cook learns when to close the book, scribble 'stock: done, in fridge, 2L' on a sticky note, and move on. That's the mechanism here: AutoCompact trains a coding agent to decide when its earlier exploration is stale, summarize the working state worth preserving, and continue from a compressed context — not as a bolt-on heuristic but as a learned part of the agent's own policy. The committed claim: coding agents can be trained end-to-end — through supervised fine-tuning followed by reinforcement learning with task-success rewards — to make compaction decisions that improve downstream task completion. This is not a new prompting trick or a fixed-window summarizer. The agent learns WHEN to compact, WHAT to keep, and HOW to continue, all jointly optimized with its coding ability. The result is a +9.2% absolute improvement on SWE-bench Verified and +5.0% on SWE-PolyBench Verified over the base model. The training pipeline is the interesting engineering. They run a base agent on coding tasks, then use a judge model to review compaction decisions, summaries, and post-compaction actions. When the judge finds flaws, corrected outputs replace the originals and the trajectory continues from the corrected decision — a form of online DAgger-style data collection where you never train on garbage rollouts. This corrected-trajectory dataset feeds supervised fine-tuning, and then RL with binary task-success rewards jointly sharpens both coding and compaction. Where does this sit on the ladder? The baseline is the unmodified base coding agent (likely a fine-tuned LLM in the SWE-agent family). The +9.2% gain on SWE-bench Verified is substantial — SWE-bench is the community's standard for repository-level coding, and single-digit absolute gains are meaningful in a regime where top systems cluster between 30-50% pass rates. The paper does not compare against alternative context-management strategies like RAG-based retrieval or sliding-window approaches in a controlled ablation, which limits how much we can attribute to the learned compaction vs. the additional SFT+RL training signal. The architecture sits in the policy-learning family for LLM agents: SFT from corrected demonstrations → RL from environment reward. The compaction mechanism itself is a special action in the agent's action space — the model outputs a compact/don't-compact decision plus a summary when it compacts. The key hardware property is that this works across inference budgets: the gains hold with a 256K context window (where overflow never happens, so compaction is purely about quality) and with a 16K window (where overflow triggers fallback compaction). This is important — it means compaction helps even when you aren't forced to compact, suggesting the agent genuinely learns to shed noise. Integrity is decent but has gaps. SWE-bench Verified is the gold-standard community benchmark for this task, and SWE-PolyBench adds cross-language coverage. But the judge model that generates corrected trajectories is not independently validated — we're trusting the judge's quality without a separate human evaluation of judge accuracy. The RL reward is binary task success, which is clean and hard to game, but we don't see ablations isolating how much of the gain comes from the compaction decisions themselves vs. the additional RL training on coding behavior. The milestone that matters: if learned context management at 16K matches or beats naive 256K performance, that's a compute-efficiency unlock — smaller context windows mean faster inference and lower cost per task. The paper shows gains at both window sizes but doesn't directly compare 16K-with-compaction vs. 256K-without. That comparison is the next concrete number the field needs. The obvious successor experiment the authors didn't run: applying AutoCompact to non-coding agentic tasks (web browsing, research assistants, multi-step reasoning) to test whether learned compaction generalizes beyond the SWE-bench domain. My read: they're saving it for the next paper — the method is domain-agnostic in principle, and the SWE-bench results are strong enough to publish standalone.