Imagine watching a puppet show where the puppeteer is really good at moving the mouth but forgot the puppet's head should also nod and tilt when it speaks. You wouldn't notice the mouth — you'd notice the head is weirdly still. That's the core mechanism here: real speech creates a coupled motion between your lips and your head pose, and every LipSync forgery method breaks that coupling in its own characteristic way. The committed claim: LipDA is the first framework to jointly detect AND attribute LipSync forgeries by exploiting the inconsistency between lip motion and head pose dynamics. Detection tells you it's fake; attribution tells you which generator made it. Prior deepfake detectors focused on pixel-level artifacts (blending boundaries, texture inconsistencies) that modern LipSync methods have learned to eliminate. LipDA shifts the battleground to temporal biomechanics — a domain where forgers have to fight physics, not just pixel statistics. The architecture is a dual-stream contrastive learning setup. One stream encodes lip motion features; the other encodes head pose features. For detection, the model learns that in real videos these streams are tightly correlated and in fakes they diverge — a contrastive loss quantifies the discrepancy. For attribution, the model captures unique temporal dynamics and audio-visual sync patterns that act as generator fingerprints. This is a supervised learning approach on video features, not a generative model or an adversarial setup. The key structural insight is treating lip-head coupling as a measurable biological signal rather than trying to spot visual artifacts. The ladder position is strong. The paper reports over 97% AUC for detection and 97.5% accuracy for model attribution, tested across two existing challenging LipSync datasets plus a new large-scale multi-generator dataset (LipSync-A) the authors built themselves. The paper claims significant outperformance over existing methods, though the abstract doesn't name specific baselines or margin sizes. The ICML 2026 acceptance suggests peer reviewers found the comparisons credible, but without seeing the full tables we're taking the headline numbers on some trust. Integrity has both strengths and gaps. On the strong side: code and the new dataset are publicly released, the method is tested on multiple datasets including one with multiple generators, and ICML peer review is a meaningful filter. On the weaker side: two of the test datasets are the authors' own construction (though this is standard when existing benchmarks are inadequate), and there's no independent replication yet. The multi-generator evaluation is important because single-generator detection is a much easier problem — the fact that they built LipSync-A specifically to test cross-generator robustness shows methodological care. The milestone that matters: current LipSync detection works on known generators. The real frontier is zero-shot detection — catching forgeries from generators the detector has never seen. LipDA's biomechanical approach should theoretically generalize better than artifact-based methods because the lip-head coupling is generator-agnostic, but the abstract doesn't report zero-shot results explicitly. If this approach maintains >90% AUC on unseen generators, that's the number that moves the field from academic benchmarks to deployable forensics. The obvious next experiment they didn't run (or didn't report in the abstract): adversarial robustness testing where a generator is explicitly trained to preserve lip-head coupling. If an adversary knows you're looking at head-pose consistency, they can add a head-pose matching loss to their generator. The authors likely know this — it's the standard cat-and-mouse dynamic in forensics — and my read is they're either saving it for follow-up work or the results were mixed enough to omit from the headline claims.