Imagine you're a building security guard watching a hundred camera feeds. You can SEE a suspicious person from the moment they walk in — they're on every screen from the lobby onward. But the decision to actually unlock the restricted door happens in one specific room, deep inside the building, in the last few hallways before the vault. That's what this paper finds inside large language models: the attacker's signal is readable from layer one, but the model's decision to comply with the attack concentrates in a narrow bottleneck in the final third of the network. The committed claim: prompt injection compliance in LLMs is not distributed diffusely across the network but is causally localized to a compact linear subspace in late layers, and patching that subspace reverses compliance in 77-92% of cases across models from 4B to 32B parameters. This is not a detection paper or a benchmark paper — it's a mechanistic dissection that identifies WHERE the decision happens, not just HOW OFTEN it happens. The method is layer-by-layer causal activation patching — the same technique that has become the workhorse of mechanistic interpretability since Meng et al.'s work on factual recall. The authors run a clean prompt (no injection) and a corrupted prompt (with injection) through the model, then systematically swap activations layer-by-layer to measure which layers causally control whether the model complies. The key finding is a clean dissociation: linear decodability of attack information is high from layer one (the model 'sees' the injection immediately), but causal leverage — the ability to flip the model's behavior by intervening — doesn't peak until the final third. The compliance mechanism occupies rank-8 subspaces in 4B and 14B models, scaling to rank-64 at 32B. That's remarkably compact. The ladder here matters. Prior work on prompt injection has been overwhelmingly behavioral: measuring attack success rates, testing defenses, cataloging failure modes. Mechanistic work on LLM internals (Anthropic's circuit-level analyses, Meng et al.'s ROME/MEMIT on factual recall) has explored editing and localization, but applying causal patching specifically to the prompt injection compliance decision is genuinely new territory. The closest predecessor is probably the growing body of work on 'refusal' circuits and safety-relevant representations, but those focus on RLHF-trained refusal behavior, not the system-prompt-override dynamic specific to injection. The integrity picture is solid but bounded. The study covers five models across 4B-32B parameters and multiple model families, which gives real architectural generality. The convergence of causal leverage and detection performance at the same layer is the strongest piece of evidence — it's hard to dismiss the bottleneck as an artifact when it also turns out to be the optimal detection site. The obfuscation test (leetspeak substitution degrading early-layer classifiers but not late-layer ones) is a smart robustness check. But the validation is still self-contained: no independent replication, no adversarial red-teaming of the detection approach by outside teams, and the obfuscation attacks tested are relatively simple. The practical implication is direct and significant: if compliance lives in a rank-8-to-64 subspace, you can build targeted interventions — either detection probes or activation-steering defenses — at a specific network location rather than bolting on external classifiers or prompt-engineering your way to safety. The 77-92% compliance reversal rate from patching is striking because it suggests the mechanism is not only localized but compact enough to be manipulable. The scaling from rank-8 at 4B to rank-64 at 32B raises the obvious question of what happens at 70B+ and frontier-scale models. The successor experiment the authors didn't run is the one everyone will ask about: does this hold at 70B+ or frontier scale (GPT-4-class, Claude-class), and can the identified subspace be used for a real-time defense that doesn't degrade normal performance? The honest read is (a) compute budget — causal patching at 70B+ is expensive — combined with (c) this is clearly the next paper. The rank-scaling trend (8 → 8 → 64 as model size grows 4B → 14B → 32B) is suggestive but needs more data points to know if the subspace stays tractably small at frontier scale.