Imagine you're a teacher grading group projects, but instead of reading the papers yourself, you ask each student to grade their own contribution. Two students collude: they submit garbage work but report perfect scores, and because you weight contributions by self-reported quality, their garbage gets amplified into the final product. That's the attack surface this paper exposes in federated learning for UAV GPS spoofing detection — and then it builds the equivalent of actually reading the papers yourself. The committed claim: federated UAV spoofing detectors that weight clients by self-reported validation accuracy are exploitable, and replacing self-reports with coordinator-side counterfactual probing — testing what each model does on synthetic spoofed samples rather than what it says — eliminates this attack vector. Two compromised clients out of ten, using a combination of data poisoning, update scaling, and inflated accuracy reporting, push backdoor lift to +0.3036. With receiver-domain behavioral probing, that collapses to -0.0265, statistically indistinguishable from zero. The mechanism is elegant in its simplicity. The coordinator constructs counterfactual spoofed samples by driving each discriminative GPS feature to its benign value, then evaluates every submitted model against these probes. Honest models behave consistently; poisoned models reveal themselves through behavioral divergence on these synthetic edge cases. This sidesteps the core weakness of Byzantine-robust aggregation methods like Multi-Krum, which require knowing the exact attacker count. When Multi-Krum's configured attacker count is wrong, it degrades sharply from +0.0061 to +0.2837 lift. The probing approach needs only an honest majority. Architecturally, this is a defense-layer paper, not a new ML architecture paper. The underlying spoofing detector is a standard supervised classifier over GPS receiver features; the contribution is the aggregation-time validation mechanism layered on top of any federated learning setup. The counterfactual sample generation leans on domain knowledge of GPS signal features — you need to know which features are discriminative for spoofing to construct meaningful probes. The integrity picture is mixed but honest. Evaluation uses a single public single-receiver dataset partitioned into simulated IID clients — not a real multi-UAV deployment, not heterogeneous data distributions. The authors explicitly report where the mechanism fails: under strong client heterogeneity, which is exactly the condition you'd expect in a real fleet with different flight paths, environments, and receiver hardware. This is the kind of limitation disclosure that builds trust, even as it narrows the claim. The scaling question is the obvious next step. Ten simulated clients on one partitioned dataset is a proof of concept, not a deployment validation. Real UAV fleets face non-IID data distributions by default — different aircraft see different GPS environments. The authors flag this weakness but don't attempt the non-IID experiment. My read: this is likely a compute-and-data limitation (they'd need multi-receiver datasets that don't exist publicly) rather than a hidden negative result. The mechanism's reliance on honest-majority assumptions also needs stress-testing at higher compromise ratios. The broader field question this participates in — how to secure federated learning against Byzantine and backdoor attacks without requiring global knowledge of the threat model — is one of the live fights in FL security. Most defenses (Multi-Krum, trimmed mean, FoolsGold) assume you know something about the attack shape. This paper's contribution is shifting validation from self-report to behavioral observation, which is a cleaner trust model. Whether it survives contact with real heterogeneity is the open question.