Imagine you're running a long meeting and instead of trusting a secretary to take notes, every participant can edit the shared document in real time — deleting irrelevant points, restructuring priorities, adding summaries. Now imagine that over many meetings, each participant learns what note-taking strategies lead to the best decisions. That's the core mechanism of Context Language Models: hand the context window back to the model as a mutable file, and let the model learn what to keep, discard, and reorganize. The committed claim is architectural, not incremental: rather than bolting context-management strategies onto a model via external scaffolding (RAG pipelines, summarization chains, sliding windows controlled by an orchestrator), CLMs treat the context as a first-class editable artifact. The model reads from and writes to a file that IS its context. This is a genuine design-level shift, not a prompting trick. The zero-shot version — using existing models with no additional training — already outperforms current SOTA context management across three distinct benchmarks: BrowseComp-Plus (11.4% accuracy gain, 21.5% fewer FLOPs), 12-hour EdgeBench (5% score gain, 59% fewer FLOPs), and a 24-hour multi-repo agent-swarm task (65% greater improvement at matched compute). The ladder here is measured in two currencies simultaneously: accuracy AND compute. Most context-management strategies trade one for the other — you summarize aggressively and lose detail, or you keep everything and burn FLOPs. CLMs improve on both axes concurrently, which is the harder result. The baselines are described as 'SOTA context management strategies,' though the paper names specific benchmarks rather than specific competing methods by architecture. The most interesting ladder result is the RL-trained variant: Qwen3.5-9B on BrowseComp-Plus improves 47.6% while using 12% fewer FLOPs. That's a Pareto improvement, not a tradeoff. Architecturally, CLMs sit in the autoregressive transformer family but add a read-write file interface between inference steps. The key structural choice is treating context as a persistent mutable state rather than an append-only log. This naturally extends to multi-agent systems where multiple agents maintain separate context files — a property the paper explicitly demonstrates on the 24-hour swarm task. The serving optimization (Suffix Cache Reuse) exploits the fact that CLM contexts share long common prefixes between edits, reducing server-side compute by 35% relative to standard SGLang. Integrity is mixed. The benchmarks are real community tasks (BrowseComp-Plus, EdgeBench), not synthetic setups designed to flatter the method. The compute measurements are reported in FLOPs, which is the right unit. However, the paper comes from a large team spanning UW, AI2, and Meta — groups with strong incentives and resources to present favorable comparisons. No pre-registration is mentioned. The RL training loop uses natural-language instruction evolution, which introduces degrees of freedom in how the optimization target is specified. The 35.9-point accuracy gain from skill optimization on a 'context-management task' is striking but the held-out evaluation details matter enormously. The milestone question is about agent autonomy at scale. Today's result: models managing their own context over 24-hour multi-agent tasks with measurable accuracy and efficiency gains. The next concrete threshold is whether CLMs can maintain coherent context management over week-long or month-long agent deployments — the kind of persistent operation that would make autonomous software engineering agents viable. The gap between 24 hours and 168 hours is not just 7x in time; it's an order of magnitude in context drift, error accumulation, and strategy brittleness. The obvious experiment not run: training the RL context-management strategy on the multi-agent swarm task itself, rather than demonstrating RL on single-agent BrowseComp-Plus and zero-shot CLMs on the swarm. The multi-agent setting is where learned context management would matter most — coordinating what to keep across agent boundaries — but RL in multi-agent settings is notoriously expensive and unstable. The honest read is (a): they ran out of compute budget for multi-agent RL, and this is the next paper.