Imagine you're an archivist at a law firm. Every time a new contract revision arrives, your predecessor's system shreds the old version into index cards — a few keywords, a summary sentence, maybe a knowledge-graph edge. When a partner asks "What was our liability cap in the March draft?", you're stuck reconstructing from fragments that were never designed for that question. The smarter move: keep every draft whole, stamped with its date, and only pull the relevant versions when someone asks. That's the core mechanism of Mem++. Where most LLM memory systems — MemGPT, A-Mem, Zep, HippoRAG — compress documents into distilled artifacts at ingestion time (facts, notes, graph edges), Mem++ does nothing generative at write time. It stores each document verbatim with metadata (date, author) and defers all intelligence to read time. When a query arrives, a temporal filter selects only documents dated up to the moment the question references, then a hybrid lexical-semantic ranker (BM25 fused with embedding similarity via reciprocal rank fusion) surfaces the relevant subset for the answering LLM. The claim is straightforward but consequential for the organizational-memory subfield: write-time distillation destroys information you'll need later, especially when questions are temporal ("What was the policy in Q2?") or versioned ("Which draft changed the delivery clause?"). By keeping everything and filtering at read time, you preserve optionality. The results on OrgMemBench are strong: Mem++ beats the best memory-system baseline (A-Mem) by 8.0 points with GPT-4.1-mini and 13.1 points with GPT-4.1-nano. With GPT-4.1-mini, Mem++ also edges out vanilla RAG by 2.6 points overall. On LoCoMo it achieves the best average LLM-judge score, and on LongMemEval-S it ranks second only behind its own entity-graph variant. Architecturally, this is a retrieval-augmented generation system with a deliberate anti-pattern: no write-time processing. The retrieval pipeline is BM25 + dense embedding with reciprocal rank fusion — well-understood, nothing exotic. The novelty is entirely in the design philosophy: treat documents as immutable, time-stamped records and push all reasoning to the reader. The entity-graph variant adds a lightweight knowledge graph for entity lookups but doesn't replace the core non-destructive store. Integrity is mixed. OrgMemBench is the primary evaluation target and appears designed to test exactly the temporal-versioning failure mode Mem++ addresses, which is fair but also favorable terrain. The authors also evaluate on LoCoMo and LongMemEval-S, both community benchmarks, which adds credibility. Code is released on GitHub. However, the baselines compared — MemGPT, A-Mem, Zep, HippoRAG — are all memory-system baselines; the RAG baseline is vanilla. There's no comparison against more sophisticated RAG pipelines with reranking, date-aware chunking, or multi-hop retrieval that production systems actually use. The answering models are GPT-4.1-mini and GPT-4.1-nano — no open-source models, which limits reproducibility outside the OpenAI API ecosystem. The real question this paper surfaces is whether the LLM memory field has been solving the wrong problem. If write-time compression is a lossy bottleneck, then the entire architecture of systems like MemGPT (which maintains a compressed working memory) is pointed in the wrong direction for versioned, multi-author organizational contexts. But this claim has a scope limit: Mem++ works because organizational documents are discrete, dated, and bounded. For streams of unstructured conversation or unbounded personal memory, raw storage without compression may not scale. The paper doesn't test that boundary. The successor experiment is obvious: scale testing. What happens when the document store grows to thousands or tens of thousands of documents spanning years? BM25 + embedding retrieval is fast, but the answering LLM's context window becomes the bottleneck. The authors likely know this and are either working on hierarchical retrieval or saving it for a follow-up. The other missing piece is a comparison against production-grade RAG with temporal metadata — the kind of system a company would actually deploy today.