Imagine you're lost in a forest and someone hands you a GPS coordinate for 'the center of the nearest town.' Useful, right? You walk toward it. But now imagine someone asks you to determine which side of a fence you're on — north or south — and all you have is that same town-center coordinate. Knowing where the town is tells you nothing about where the fence is or which direction crosses it. That's the core mechanism this paper dissects: a safety detector that knows where 'safe' is located in embedding space, but has no idea which direction leads to 'unsafe.' The committed claim: the 'positive centroid rule' — scoring response safety by cosine similarity to the mean embedding of known-safe responses — does not reliably separate safe from unsafe text. On two human-labeled corpora (PKU-SafeRLHF and Aegis) across four frozen encoders, the raw safe-prototype score hits ROC-AUC of 0.457–0.545, with two experimental cells significantly below chance. That's not just bad; it's sometimes inverted. When the authors switch to an explicit safe-minus-unsafe reference direction, performance jumps to 0.588–0.738 on the same embeddings and the same data. On a jury-labeled control set (BeaverTails), the prototype inverts entirely (0.358–0.405) while the reference reaches 0.754–0.793. The architecture story is simple and that's the point. This isn't proposing a new model. It's auditing an existing scoring rule — specifically the one used in a recent sleeper-agent detector — by running it across four off-the-shelf frozen encoders (the paper doesn't fine-tune anything) on prompt-controlled, human-labeled splits. The method under audit belongs to the one-class classification family: define the positive class, measure distance to it, threshold. The paper argues this family needs a reference direction (safe minus unsafe), not just a reference location (safe centroid), to work. The integrity design is notably careful for a 15-page workshop paper. Prompt-grouped splits prevent prompt-label confounding from inflating scores — and a dedicated ablation demonstrates that prompt-only signals can do exactly that in uncontrolled evaluations. The authors test across four encoders rather than picking the one that tells the best story. They report the cells where their own referenced method doesn't win (BeaverTails at the 5% false-safe threshold). The scoping is explicit and honest: this is about the raw positive centroid, not all one-class methods or safety-specialized guards. The practical fallout is immediate. Anyone deploying a sleeper-agent detector or safety filter that scores by proximity to a safe centroid should worry that their detector may be operating at or below chance on real distributions. The fix isn't exotic: 80–634 labeled unsafe responses are enough to recover most of the referenced direction's ranking power. A fully unlabeled held-out reference recovers partial-to-most of it, but degrades sharply when only 5% of the pool is unsafe — a realistic production scenario. The paper's scope is deliberately narrow, and that's a strength. It doesn't claim all embedding-based safety methods are broken. It shows that one specific, recently proposed scoring rule — cosine similarity to the safe centroid — lacks the geometric information needed to orient a decision boundary. A class mean is a location, not a direction. The distinction sounds pedantic until your detector is flipping coins on real data. What's left on the table is scaling this audit to safety-specialized encoders (the paper uses general-purpose frozen ones) and to production-scale distributions where the unsafe base rate is far below 5%. The authors don't run those experiments, and the honest read is resource constraints plus scope discipline for a workshop paper — not that they tried and it failed.