Imagine you're moving into a new apartment. You can either keep everything in boxes stacked by the front door (easy to grab, but the hallway fills up fast and you trip over them daily), or you can unpack items into drawers and shelves where they belong (slower upfront, but the hallway stays clear and you find things by muscle memory). This paper is about the second strategy for LLMs: instead of stuffing new knowledge into the context window — the hallway — you bake it into the model's parameters — the drawers. The committed claim: this is not a result paper. It is a survey that proposes a 2D taxonomy for organizing all methods that write reusable knowledge into parameter-like objects (adapters, soft prompts, modified FFN weights) rather than into the context window. The two axes are Parameter Placement (where in the architecture the memory lives: embedding, attention, FFN, or hybrid) and Parameter Acquisition Time (when the memory was created: offline before deployment, or online during it). The authors argue this grid clarifies a landscape that the field has been navigating without a shared map. The ladder question is unusual for a survey. There is no single baseline to beat. The implicit comparison is against in-context learning (ICL), and the survey's thesis is that parametric memory is complementary to ICL: it eliminates the per-inference cost of re-encoding long contexts, trades context capacity for parameter storage, and makes memory reusable across calls. The paper cites concrete methods — LoRA adapters, knowledge neurons, function vectors, soft prompts — but does not run head-to-head experiments itself. It organizes others' numbers rather than producing its own. Architecturally, the survey covers a broad family tree: soft-prompt and prefix-tuning methods (embedding layer), attention steering via function vectors and task vectors (attention layer), knowledge-editing and LoRA-style methods (FFN layer), and hybrid approaches that touch multiple layers. The key structural insight is that placement determines what kind of knowledge is encoded: embedding-layer methods handle style and task framing well but struggle with factual recall; FFN-layer methods excel at factual updates but can catastrophically interfere with existing knowledge. The online/offline axis maps neatly onto the deployment engineering question of whether you can afford training-time compute or need to learn on the fly. Integrity is the survey's weakest dimension. The authors synthesize existing literature and propose a framework, but they do not run any new experiments validating that their taxonomy predicts real behavior (e.g., that placement determines interference patterns). The framework is plausible and well-organized, but it remains an editorial contribution rather than an empirical one. The open-directions section on interference, safety, ICL co-design, and recursive self-improvement is thoughtful but does not provide quantitative grounding. The milestone this survey implicitly tracks is the convergence of parametric memory with agentic LLM systems. The authors frame the endgame as models that continuously learn from interaction and store that learning in composable parameter objects — essentially giving LLM agents long-term memory that doesn't consume context. The concrete next number to watch is whether any online parametric memory method can match RAG retrieval quality on standard QA benchmarks while using zero context tokens for the retrieved knowledge. The obvious experiment not run: a controlled head-to-head benchmark comparing the four placement categories (embedding, attention, FFN, hybrid) on the same tasks with the same base model. The honest read is (a) — this would require substantial compute to do fairly across all categories, and a survey paper from a university group likely didn't have the budget. It's the natural next paper, and the taxonomy they've built is essentially the experimental design waiting to be executed.