Imagine you're a paralegal reviewing a massive discovery dump — thousands of documents, arriving one folder at a time. You can't keep everything on your desk (the desk is fixed-size), so you have to decide what to file permanently, what to toss, and what to leave in your inbox for now because you're not sure yet if it matters. The worst mistake isn't filing the wrong thing — it's shredding a document before you realize the next folder makes it critical. That's exactly the problem CoEM solves for LLMs doing long-context reasoning. The committed claim: a learned three-way policy (promote, retain, or discard) over a pending set of verbatim source excerpts, trained with reinforcement learning, consistently outperforms fixed-compression memory baselines on long-context QA. On 6,400-document inputs using Qwen3.5-9B, CoEM beats the strongest memory baseline by 10.4–11.4 F1 points. This is not a new architecture — it's a new memory management regime bolted onto existing LLMs. The mechanism is straightforward but well-constructed. Under a fixed context-memory budget, CoEM maintains two pools: a pending set of verbatim source excerpts and a committed memory of compressed facts. As new chunks arrive, a learned policy revisits each pending item and makes a three-way decision: promote it to committed memory (compress it now), keep it pending (wait for more context), or discard it. A frozen verifier gates promotion — proposed memory facts must be grounded in retained excerpts and current context. The whole system trains via RL combining step-level evidence rewards with final-answer rewards. Where does this sit on the ladder? The paper compares against ReadAgent, MemWalker, and several chunk-and-compress baselines. The +10.4–11.4 F1 gain over the strongest baseline is substantial, and importantly the experiments test at genuinely long contexts (6,400 documents). However, all baselines are memory-augmented reading strategies, not pure long-context models with extended windows (like Gemini 1.5 Pro at 1M tokens). The paper acknowledges this scope — it's competing within the bounded-memory paradigm, not claiming superiority over unlimited-context approaches. The integrity picture is mixed-positive. The evaluation uses established QA benchmarks (HotpotQA, MuSiQue, 2WikiMultihopQA) with public test sets, which is good. The ablation study is thorough — they isolate the pending mechanism, the RL training, the verifier, and the reward shaping. But the evaluation is same-team, single-model-family (Qwen), and there's no pre-registration. The 6,400-document scale is impressive but synthetic in the sense that documents are concatenated to create long contexts. The real contribution here is conceptual: premature compression is a named, quantified failure mode, and deferred commitment is a trainable solution. The pending-set mechanism is model-agnostic in principle — it could sit atop any chunk-processing LLM pipeline. The RL training recipe (step-level evidence rewards + final-answer rewards) is the kind of specific, reproducible detail that makes follow-on work possible. The obvious next experiment is scaling this to truly heterogeneous, real-world corpora (legal discovery, medical records, codebases) where the distribution of when evidence becomes relevant is far less predictable than in multi-hop QA benchmarks. The authors likely didn't run this because the RL training loop is expensive and QA benchmarks are the accepted evaluation currency. The more interesting missing experiment: testing against extended-context models directly to see if smart memory management on a 32K window can match a 128K or 1M native window.