Imagine you're renovating a house with three different contractors. Each one hands you their final blueprint — but every blueprint includes the original house's layout baked in. If you try to overlay all three, you get triple-counted walls and conflicting foundations. The fix is obvious once you see it: ask each contractor for only the diff — what they changed from the original plan. That's Δ-MOPD. The committed claim: when distilling from multiple teacher models into a single student, transferring each teacher's logit shift (teacher minus its base) re-anchored at the student's own base outperforms transferring the teacher's raw endpoint policy — specifically when multiple teacher signals are combined at the same state. The gain is +4.11 on math and +1.95 across five benchmarks with three composed teachers. With two teachers, it merely matches endpoint accuracy. The mechanism is clean and well-diagnosed. The authors show that in endpoint transfer, the 'inherited base pull' — preferences baked into the teacher's pre-training base that have nothing to do with post-training alignment — can actually exceed the magnitude of the post-training shift itself. When you combine multiple teachers, you're stacking multiple copies of this irrelevant base signal. Subtracting it out reduces the teacher-term norm ratio and the KL divergence between the composite target and the student. This is not a learned trick; it's an algebraic decomposition. Where it sits on the ladder: the paper compares Δ-MOPD against standard endpoint multi-teacher distillation, holding teacher selection fixed. This is the right control — isolating target construction from teacher choice. The baselines are not stale; on-policy multi-teacher distillation is the current production pattern at scale (Meta is a co-affiliation). The gains are modest but consistent: meaningful in the composition setting, neutral in interleaved routing where only one teacher speaks at a time (which makes sense — no base-pull stacking when signals aren't combined). The integrity picture is mixed. The experimental setup is careful — two settings (composition vs. routing), multiple teacher counts, both phase orders tested for routing — and the authors are honest that two-teacher composition merely ties. But we don't see dataset details, model scale, or compute budgets. There's no code release mentioned, no pre-registration, and the benchmarks appear chosen post-hoc. The phased-routing result showing order-gap reduction from 10.50 to 6.42 is interesting but comes with the caveat 'supporting evidence that the benefit may extend,' which is appropriately hedged. The practical upshot: if you're running multi-teacher on-policy distillation in production — composing reward signals from safety, helpfulness, and coding teachers into one student — this is a free architectural improvement. The logit subtraction costs nothing at inference and the implementation is trivial. The paper frames target construction as an independent design axis from teacher selection, which is the right abstraction and suggests a combinatorial space of further improvements. What's missing is scale. The authors don't report model sizes, don't test at frontier scale, and don't compare against other composition strategies like weighted averaging of teacher outputs or mixture-of-experts routing. The obvious next experiment — scaling to 70B+ students with 5+ teachers — would tell us whether the base-pull problem gets worse or better with scale. My read: they're saving it for the next paper, likely with production-scale results at Meta.