You know how Track Changes in Word can embarrass you — a lawyer sends opposing counsel a "clean" document, but the metadata still contains every red-lined insult from the partner review? LLMs have invented their own version of this problem, except it's baked into how they communicate, not into a file format. When a user says "remove the password before sharing," the model dutifully deletes the password from the config file — then adds a helpful comment: "Removed the password 'No4!' as requested." The recipient, who was never supposed to see the password, reads it in the comment. The paper calls these "revision traces," and they're disturbingly common. The committed claim: LLM assistants systematically leak withdrawn information through their own narration of edits, and this is a new, uncharacterized class of information disclosure. The authors back this with both in-the-wild evidence and controlled experiments. Across three public conversation corpora (ShareGPT, WildChat, LMSYS-Chat-1M), they find 26,753 revision requests, of which 8.8% (2,363) leave revision traces — meaning the model's response explicitly references the thing the user asked to remove. That's not a lab curiosity; that's real users leaking real data in the wild. To study the problem under controlled conditions, they introduce RevLeakBench: 100 tasks across five scenarios (personal data, credentials, internal strategy, medical records, legal details), each with a conversation track and an agent track. Across six models (GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro, Llama 3.1, Qwen 2.5, DeepSeek-V3), roughly half of deliverables mention the edit after a revocation request. More critically, a third-party reader who sees only the final deliverable can recover the withdrawn item from about 13% of them. Even when the model is explicitly told "your entire reply will be forwarded to the recipient," revision traces persist in 36.4% of deliverables. The models are constitutionally chatty about their own editing process, and awareness of an audience only partially suppresses the behavior. Architecturally, this is a behavioral vulnerability rooted in RLHF-trained helpfulness norms, not a classical injection or side-channel attack. The models aren't being tricked — they're doing what they were trained to do (explain changes, be transparent about edits), but in a context where that transparency is the threat. The paper tests prompt-level defenses ("do not mention removed items") and finds them partially effective but leaky. Their best mitigation is an output-side filter — a post-generation scan that catches revision traces before delivery — which sharply reduces recovery rates with minimal loss of required content. The integrity setup is solid for a first-characterization paper. The in-the-wild analysis uses three independent public corpora, not synthetic data. The benchmark is new (RevLeakBench), so there's no cherry-picking against existing leaderboards, but there's also no pre-registration — the benchmark was designed alongside the study. Six commercial and open-source models are tested, which is a reasonable spread. The main weakness is that "recovery rate" depends on the judgment of an evaluator (human or LLM-as-judge), and the boundary between a suggestive trace and a recoverable item is subjective. The milestone question is practical: what recovery rate is acceptable? The authors get it down substantially with their output filter, but don't specify a target. The real unlock is whether major LLM providers adopt output-side filtering as a default for agentic and drafting workflows. If this class of leak becomes part of standard red-teaming, the paper will have moved the needle. If it stays an academic curiosity, the 13% recovery rate will persist quietly in every enterprise deployment that uses LLMs for document preparation. The obvious experiment not run: adversarial prompting to maximize trace leakage. The paper studies the default case — normal users making normal revision requests — but doesn't test whether a malicious third party who controls part of the conversation context can engineer prompts that increase trace rates. The honest read is this is being saved for a follow-up paper, since it opens an entirely different threat model (adversarial vs. accidental disclosure).