Imagine you're a doctor in a busy ER. A patient arrives with chest pain and you have their EKG, blood work, and medical history — you can diagnose confidently. Another patient arrives with the same complaint but you only have a blurry photocopy of a referral letter from three years ago. A bad system gives both patients the same confidence score. A good system says: "I think this is X, but I don't have enough information to commit." SAFE-MR is that second system, applied to multimodal rumor detection. The committed claim: SAFE-MR is the first framework to explicitly decouple evidence sufficiency from veracity prediction in multimodal misinformation detection, using separate prediction heads and evidence-perturbation training to teach the model when to abstain. This matters because most evidence-augmented detectors treat retrieved evidence as uniformly informative — if you found five web results, you're good to go. In reality, those five results might all be duplicated reports from the same source, or they might contradict each other without resolution, or they might simply not address the claim's core assertion. SAFE-MR forces the model to reckon with this gap. The architecture works in three stages. First, it decomposes image-text social media posts into discrete verifiable claims — breaking a complex post into atomic checkable units. Second, it builds a relation-aware claim-evidence graph that encodes not just whether evidence exists but its provenance and contextual compatibility with each claim. Third, separate "veracity" and "sufficiency" prediction heads allow the system to output both a truth judgment and a confidence-in-evidence score. The sufficiency head enables selective prediction: when evidence is insufficient, the system can abstain rather than guess. The training regime is where the real engineering insight lives. Evidence interventions — deliberately adding irrelevant evidence and removing relevant evidence during training — teach the model two properties simultaneously: stability (don't change your mind when noise is added) and sensitivity (do change your mind when real evidence is removed). This is a form of causal regularization, forcing the model to learn which evidence actually supports its predictions rather than pattern-matching on evidence volume. On three benchmarks — NewsCLIPpings, VERITE, and XFacta — SAFE-MR achieves macro-F1 scores of 91.2%, 75.8%, and 85.2%. The more telling numbers are the gains over the matched backbone with evidence: +2.2, +4.9, and +4.8 percentage points respectively. These aren't world-shattering deltas, but they're consistent across datasets with very different characteristics (manipulated images, out-of-context media, cross-lingual claims). The selective prediction results are stronger: AURC drops from 0.105 to 0.075 and error at 80% coverage falls from 13.8% to 8.5%, meaning the system makes substantially fewer mistakes when it's allowed to say "I'm not sure." The honest limitation: all three benchmarks are research datasets with curated evidence. The gap between curated evidence retrieval and real-world evidence retrieval — where search engines return SEO-optimized junk, paywalled sources, and adversarially manipulated content — is enormous. The sufficiency head is only as good as its training distribution, and that distribution does not yet include the adversarial mess of actual fact-checking at scale. The deeper contribution here isn't the F1 numbers — it's the framing. Most misinformation detection research treats the problem as pure classification: rumor or not-rumor. SAFE-MR argues that the epistemically responsible question is two-dimensional: what do I believe AND how well-supported is that belief? This is a framework contribution that could propagate beyond rumor detection into any evidence-augmented decision system.