Imagine you suspect someone photocopied your exam answers. You can't see their copying process, but you CAN compare their answer sheet to yours and ask: are these suspiciously similar in the specific ways that copying produces, versus the ways that two smart students independently solving the same problems would overlap? That's the core mechanism here — not checking IF the answers match, but checking whether the PATTERN of matching looks like distillation versus independent learning. The committed claim: you can build a statistical test that distinguishes a model trained on a proprietary LLM's reasoning traces from one trained independently on just the reference answers, and you can calibrate this test using shadow models to produce real p-values. Chen and Du formulate distillation inference as a hypothesis test. The null hypothesis is that the suspect model was trained independently; the alternative is that it was distilled. Shadow models — some distilled, some independent — establish the expected behavior distribution under each hypothesis. The auditor then scores how closely the suspect model predicts the teacher's reasoning outputs and converts that score to a calibrated p-value. The experimental setup is deliberately constrained: Qwen2.5-7B as teacher, Llama-3.2-3B as the suspect architecture. The test achieves a true positive rate of 1.0 at significance level 0.02. That's a perfect detection rate with a very tight false-positive budget. But the constraint is the point — this is a poster paper establishing feasibility, not a production-ready forensic tool. One teacher model, one suspect architecture, one size pairing. The architecture is refreshingly simple. No exotic machinery — just supervised fine-tuning to produce shadow models, a scoring function measuring alignment with the teacher's reasoning traces, and standard statistical testing to convert scores into p-values. The key insight is that distilled models don't just learn the right answers — they learn the teacher's specific reasoning style, and that stylistic fingerprint is detectable. The shadow model approach is borrowed from membership inference literature, adapted here for the distillation question. On the integrity front, this is a self-contained simulation study. The authors are grading their own homework: they create the distilled models, they create the independent models, they run the test. There's no adversarial robustness evaluation — what happens when someone distills but then fine-tunes to scrub the fingerprint? The CCS'26 acceptance as a poster signals peer review found the logic sound but the scope preliminary, which is the honest assessment. The ladder question is interesting because there isn't a mature baseline to beat. Model distillation detection is nascent. Prior work on model extraction detection (watermarking, fingerprinting) addresses adjacent but distinct problems. This paper's contribution is the formulation itself — casting distillation detection as hypothesis testing with shadow-model calibration — rather than beating an established benchmark. The 1.0 TPR at α=0.02 is striking but needs stress-testing across architectures, sizes, and adversarial conditions. The milestone to watch: can this approach maintain its detection power when the suspect model is much larger (say, 70B distilled from a frontier model), when the distillation process includes intermediate fine-tuning designed to obscure provenance, and when the auditor doesn't know the exact teacher model? Those are the three escalation steps from proof-of-concept to deployable forensic tool. If the method survives even one of those, it becomes practically relevant for API providers worried about unauthorized distillation.