Imagine a student who aces every exam — but you discover they never opened the textbook. They're pattern-matching from the question alone, not reasoning from the material. That's what this paper catches LLM compliance systems doing. The models get the right answer often enough, but when you remove, swap, or negate the regulatory rule they're supposed to be applying, the verdict barely flinches. The rule is furniture, not foundation. The core claim is sharp and testable: LLM compliance verdicts are substantially invariant to perturbations of the supplied regulatory rule. The authors test this across five models (including a dedicated guard model) and 20 regulatory and platform-policy domains using two complementary metrics — Outcome Change Sensitivity (OCS), which checks if the verdict flips, and ICS-delta, which checks if the model's internal representation of compliance shifts at all. Neither moves much under rule deletion, swapping, or negation. The guard model — the one purpose-built for this job — scores 51% accuracy, barely above chance, and shows the weakest rule sensitivity of the five. The architecture of the study is diagnostic, not prescriptive. This isn't proposing a new compliance method; it's an audit framework. The experimental family is behavioral probing combined with representation-level analysis (examining internal activations, not just outputs). The key structural insight is the distinction between cases where models already "know" the answer from the scenario description alone versus cases where the rule actually provides decision-relevant information. On the latter — genuinely rule-dependent cases — models do track the rule closely. The problem is that most deployed compliance scenarios are easy enough that models coast on surface patterns. The integrity picture is mixed but honest. The authors test five named models across 20 domains with systematic perturbation types — that's a real experimental matrix, not cherry-picked examples. They also try two mitigation strategies (better prompting and direct activation intervention) and report that neither closes the gap, which is the kind of negative result that strengthens a paper. But there's no pre-registration, no independent replication, and the specific case sets are author-constructed rather than drawn from a community benchmark. The 51% guard model accuracy is devastating but comes under a "custom-rule adaptation of its native taxonomy" — meaning the evaluation conditions differ from the model's designed use case. The ladder position is unusual because this paper isn't competing on a performance metric — it's establishing a diagnostic methodology. The relevant baseline isn't "previous compliance accuracy" but "previous audits of rule-grounding." The closest prior art is adversarial robustness testing and faithfulness evaluation in NLP, but applying it specifically to regulatory compliance systems with this perturbation-based protocol is the contribution. General-purpose models hit 90-92% accuracy versus the guard model's 51%, which is the paper's most quotable number. The milestone question points toward deployment standards. Right now, compliance systems are evaluated on accuracy alone. This paper argues that's insufficient — you need rule-sensitivity metrics alongside accuracy. The practical threshold is regulatory adoption: when does a regulator require proof that a compliance system's verdict is causally grounded in the cited rule, not just statistically correlated with the right answer? That's a policy milestone, not a technical one, and it could arrive fast given the EU AI Act's transparency requirements. The obvious successor experiment is scaling to harder cases — real regulatory disputes where the rule genuinely disambiguates the outcome, not the easy cases that dominate current benchmarks. The authors show models track rules on hard cases; the question is what fraction of real-world compliance queries are hard versus easy. My read: they didn't run this because constructing a validated hard-case benchmark requires domain expert annotation at scale, which is expensive and slow. That's the (a) ran-out-of-budget explanation, and it's the honest one.