Imagine you're in a crowded office and someone shouts "Who has the Johnson file?" You know where it is—it's on Sarah's desk. But you can't just point to it yourself. You need Sarah to look up, acknowledge the file, and then pass it over. The decoding position is you; Sarah is the context token. Without her attention to the file, the information never reaches the output. That's the core mechanism this paper pins down inside Qwen3-8B, an 8-billion-parameter language model. The committed claim: entity copying—the trivially easy task of echoing a name from prompt to output—is performed by two distinct groups of late attention layers (in the second half of the model), and these layers are both necessary and sufficient. Strip away the rest of the network, and copying still works. Strip away either group, and it fails. This is not a new capability or a new architecture—it's a mechanistic dissection of a behavior every LLM already exhibits. The authors introduce two clever interventional tools. "Genie-in-a-bottle" selectively enables or disables specific layers for the copying task, giving precise layer-group ablation without retraining. "Attention lobotomy" surgically severs a token's ability to attend to entity tokens while leaving the rest of the attention distribution intact—no zeroing out rows that would distort softmax. These are well-designed causal tools, not just correlation probes. The distinction matters: they're testing necessity and sufficiency, not just activation magnitude. The surprise finding is about context tokens. The decoding position (the slot generating the answer) needs to attend to entity tokens—that's obvious. But the paper shows that context tokens flanking the entity also need to attend to the entity tokens, even though those context tokens don't themselves store entity information (unless they have special semantic properties, like being a possessive or a relational marker). Kill context-token attention to the entity, and the model can still produce a plausible answer, but it fails to copy the exact entity tokens. The context tokens function as a relay network—they don't carry the signal, but they shape the channel. Where does this sit against prior work? The mechanistic interpretability literature—Olsson et al.'s induction heads (2022), Wang et al.'s IOI circuit (2023)—has identified attention-head-level circuits for name copying in smaller models (GPT-2 scale). This paper operates at the layer-group level rather than individual heads, and on a much larger model (8B parameters), but it doesn't directly compare its findings to the IOI circuit's granularity. There's no accuracy metric to beat a baseline on; this is a structural claim about where computation lives, not a performance claim. The validation is primarily same-team interventional experiments on a single model. The integrity profile is mixed. The interventional methodology is strong—genie-in-a-bottle and attention lobotomy are causal, not correlational, and the authors test both necessity and sufficiency. But the experiments run on one model (Qwen3-8B) with no cross-model validation, no pre-registration, and no independent replication. The generality claim—that late-layer entity copying is a universal transformer property—is suggested but not tested. If these results don't replicate on Llama 3 or Gemma, the finding collapses from "how transformers copy" to "how Qwen3-8B copies." The obvious next experiment is running these same interventions on architecturally different models—Llama 3, Gemma 2, Mistral—to test whether the two-group late-layer structure is universal or Qwen-specific. The authors almost certainly know this and are either saving it for a follow-up or ran out of compute. Given that the tools (genie-in-a-bottle, attention lobotomy) are model-agnostic in design, the barrier is compute budget rather than methodological difficulty. The context-token finding also begs for a quantitative follow-up: how many context tokens need to attend, at what strength, before copying degrades?