Imagine you're running a soup kitchen but you only measure the average calories consumed across the whole neighborhood — including the people who ate dinner at home. Your average goes up a little, and you conclude the kitchen is mildly useful. You never notice that the hungriest people ate three times as much as everyone else, or that the well-fed barely touched the soup. That's what happens when AI-in-education studies report only mean effects. This paper's core mechanism is about WHERE in the distribution a treatment lands, not WHETHER it lands. The committed claim: expert-verified AI study materials (podcasts, FAQs, quiz guides) produced with a source-grounded model and checked by a named graduate teaching assistant produce a 2.34-mark advantage on a 50-mark exam component — but roughly three-quarters of that average gain originates in the bottom quintile of the grade distribution. The paper introduces a concept it calls the "judgement burden" — the expertise a learner must supply to screen AI output before learning from it — and argues that pre-release expert verification shifts this burden from students (who lack domain knowledge) to an accountable tutor. The design is a two-cohort difference-in-differences setup in a compulsory first-year economics course at a UK university: 170 students, 340 examination marks. One cohort got access to the AI-generated materials; the other didn't. The share of marks falling below the upper-second classification boundary dropped by 24.7 percentage points relative to the counterfactual. Effects were statistically significant at every threshold from 23 to 31 marks and at none above — meaning the intervention pulled up the bottom without detectably moving the top. This is a distributional finding, and it's the load-bearing result. The paper is honest about fragility in one important way: the average treatment effect of 2.34 marks is not robust to removing the lowest-scoring students from the pre-intervention cohort. Strip out the weakest performers before treatment, and the mean effect wobbles. The threshold effect — the 24.7pp drop below the boundary — survives this robustness check. This asymmetry is the paper's sharpest methodological contribution: it demonstrates concretely that mean-only reporting can mask distributional concentration. Qualitative data from 36 student interviews adds a useful mechanism. Students reported that the verification label — knowing a named GTA had checked the material — gave them a reason to engage without suspending their own scrutiny. This is a subtle but important distinction from blind trust: the verification acted as a credibility anchor, not a permission to stop thinking. The paper frames this as the difference between "ending scrutiny" and "enabling engagement," and the framing holds. Architecturally, this is not an ML paper. It's a quasi-experimental education study that happens to use AI-generated content as the treatment. The AI pipeline (source-grounded model plus human expert check) is a black box here — the paper doesn't evaluate model quality, hallucination rates, or generation methods. The contribution is entirely on the evaluation side: showing that distributional analysis is necessary and that expert verification changes the equity profile of AI-assisted learning. The limitation is scale and generalizability. One course, one university, one cohort comparison, 170 students. The economics context (structured exam, clear marking thresholds) may not transfer to essay-based disciplines or STEM lab courses. The named-GTA verification model requires institutional investment that may not scale cheaply. But the methodological argument — stop reporting only means — is field-general and immediately actionable for anyone designing AI education interventions.