Imagine you ask five different translators to highlight the most important words in a foreign-language contract. They all agree on the outcome — this contract is about a real estate sale — but each one underlines different words as the ones that clinched it. Now imagine the contract is ambiguous, maybe about a sale or maybe a lease, and the translators not only disagree on what it means but also disagree wildly on which words matter. That's the explainability disagreement problem in NLP, and this paper turns it into a measurable, auditable quantity. The committed claim: when you run five different attribution methods (SHAP, LIME, occlusion, Input×Gradient, Attention×Gradient) against DeBERTa-v3 doing zero-shot medical abstract classification, the degree to which those methods agree on which tokens matter is a direct proxy for how confident the model actually is. Strong agreement means the model is locked in; weak agreement means the prediction is fragile regardless of the confidence score. This isn't a new model or a new XAI technique — it's a comparative auditing protocol. The experimental setup is clean and deliberate. One thousand abstracts per class drawn from the Medical Abstracts corpus, classified into five diagnostic categories via natural language inference with five enriched hypotheses per category. The five explanation methods span three families: model-agnostic (SHAP, LIME), deep-learning-specific (occlusion, Input×Gradient), and transformer-specific (Attention×Gradient). Explanations are standardized to top-token attribution lists and pairwise agreement is measured via Jaccard index — a simple, interpretable metric that asks: of the tokens both methods flagged as important, how many overlap? The results land where you'd hope a careful diagnostic paper would land: unambiguous clinical categories (e.g., neoplasms, digestive diseases) produce high classification accuracy AND high inter-method agreement, while semantically overlapping categories (e.g., pathological conditions that span multiple organ systems) degrade both accuracy and explanatory consensus simultaneously. The paper identifies three specific failure mechanisms: lexical hypersensitivity (the model latches onto surface-level trigger words), semantic overlap (categories share too much vocabulary for clean separation), and loss of attribution coherence (methods diverge because the model's internal reasoning is genuinely confused, not just uncertain). The architecture choice — DeBERTa-v3 via zero-shot NLI — is interesting precisely because it's off-the-shelf. The paper isn't claiming a better classifier; it's asking what happens when you audit a strong general-purpose model with multiple explanation lenses simultaneously. The answer is that single-method explanations are dangerously incomplete. Any one of the five methods alone would give you a plausible-looking explanation that might be an artifact of that method's particular inductive bias. Integrity is mixed. The Medical Abstracts corpus is public, the methods are standard, and the Jaccard metric is transparent. But the benchmarking is self-contained — there's no comparison to any other XAI auditing framework, no pre-registration, and no code release mentioned. The balanced 1,000-per-class sampling is good practice, but the five diagnostic categories are chosen to demonstrate the framework rather than stress-test it against the full ICD ontology. The qualitative error analysis (three failure mechanisms) is insightful but not formalized into a reproducible metric. The real value here is the operational insight: if you're deploying transformer-based classifiers in clinical settings, you should never trust a single explanation method, and you should treat inter-method disagreement as a first-class diagnostic signal. The paper's recommendation to use specific clinical ontologies over broad diagnostic labels is practical and actionable. But the framework needs to be tested on harder tasks — longer documents, more categories, adversarial inputs — before it becomes a standard auditing protocol.