You know how a spell-checker can fix the word you misspelled but then start flagging every correctly-spelled word that looks vaguely similar? That's the core problem PatchBench is measuring. When you patch an LLM to stop answering a harmful prompt, the patch often bleeds into nearby territory — refusing benign prompts that happen to share structure or vocabulary with the harmful one. Aggregate metrics like attack success rate and global capability scores can't see this. PatchBench can. The committed claim: existing jailbreak repair evaluation is fundamentally incomplete because it only measures whether the patch blocks the exact harmful prompt and whether global capability is preserved. It misses the local neighborhood — the zone of prompts structurally or lexically adjacent to the patched behavior. PatchBench-Local fills this gap with a systematic protocol that generates three families of neighbors for each harmful source: harmful variants (does the patch generalize?), benign prompts with matched structure (does the patch over-refuse?), and benign prompts reusing key harmful terms (does the patch trigger on vocabulary alone?). The pipeline is industrious. Starting from 27,870 prompts drawn from 37 public datasets, the authors query 8 open-source instruction-tuned models to find prompts that actually elicit actionable harmful answers — not theoretical vulnerabilities but empirically confirmed jailbreak failures. A three-stage filter (WildGuard automated classifier, pairwise Elo ranking, manual verification) winnows this to 400 high-confidence jailbreak failures. That 1.4% yield rate tells you something important: most 'jailbreak' prompts in existing datasets don't actually produce harmful outputs from current models. The benchmark is testing real wounds, not hypothetical ones. The key experimental result is damning for current practice. Four activation steering methods were evaluated using both PatchBench-Local and MMLU. Global capability (MMLU) remained nearly unchanged — the patches look clean on aggregate. But local benign-neighbor preservation collapsed. Models patched to refuse a harmful prompt would also refuse structurally similar but completely benign prompts. This is the central finding: aggregate metrics provide a false sense of surgical precision. The patch is a shotgun, not a scalpel, and you can only see the spray pattern if you test the neighborhood. Methodologically, PatchBench sits in the activation patching family — representation engineering and steering vectors applied at inference time rather than fine-tuning. The benchmark is agnostic to the specific patching method but is designed to stress-test the precision of any intervention that modifies model behavior at the activation level. The three-neighbor protocol is the real contribution: it converts a binary question (does the patch work?) into a four-dimensional one (does it generalize to variants? does it spare benign structure-matches? does it spare benign vocabulary-matches? does global capability survive?). The integrity story is solid for a benchmarks paper. The prompt bank is derived from public datasets, the curation pipeline is documented, and the evaluation protocol is reproducible. The main limitation the authors should have pressed harder on: the 400-prompt bank covers 8 models, but the neighbor-generation quality depends on the faithfulness of the automated paraphrase and benign-variant generators. If those generators introduce systematic biases — say, benign variants that are too obviously benign — the protocol would understate collateral damage. The NeurIPS Datasets and Benchmarks acceptance is a good sign for community uptake. The real value here is the diagnostic lens, not the specific numbers. PatchBench-Local gives the field a protocol for asking: is your safety intervention precise or just aggressive? That question will outlast any particular activation steering method. The next milestone is whether this protocol becomes standard in jailbreak repair papers — if it does, the field's quality of evidence improves meaningfully. If it doesn't, we'll keep publishing patches that look clean on aggregate while silently degrading the model's willingness to answer legitimate questions.