Imagine you're training a new barista. You could hand them a stack of real customer orders annotated by your best veteran (distillation), or you could have the veteran invent fake orders from memory (synthetic data). Both sound reasonable. This paper shows that for clinical NLP, the fake orders are nearly useless—the new barista never outperforms an untrained hire unless you give them real tickets. The committed claim: a fine-tuned open-weight Gemma-3-12B model, trained on GPT-4o-labeled real radiology reports, matches GPT-4o itself at multi-label intracranial hemorrhage acuity extraction (macro-F1 0.845 vs 0.850, p=1.000), while running entirely on a single 24 GB consumer GPU. The paper frames this as a privacy-preserving, cost-stable, reproducible alternative to hosted proprietary APIs. That framing is honest—this is a practical engineering contribution, not a fundamental architecture breakthrough. The experimental design is the paper's strongest asset. A clean 2×2 factorial crosses two adaptation strategies—discriminative classification head (CH) and generative instruction fine-tuning (IFT)—with two data sources: distillation from real GPT-4o-labeled reports and GPT-4o-generated synthetic reports. Five training sizes are tested for each cell. This is how you isolate what matters. The answer is unambiguous: data source dominates. Synthetic-data models failed to exceed the un-tuned base model at any training size. The fine-tuning method (CH vs IFT) was secondary. The ladder position is clear but modest. The benchmark is GPT-4o, which is the practical SOTA for this task in clinical settings. Gemma-3-12B distilled instruction-tuned (DIFT) ties it. The base un-tuned model trails by 0.178 macro-F1. But the test set is 100 expert-adjudicated reports—adequate for a focused clinical extraction task but not a large-scale community benchmark. The ICH acuity extraction task is narrow and well-defined, which is both the strength (clinical utility) and the limitation (generalizability unknown). Architecturally, this is straightforward LoRA/QLoRA-style parameter-efficient fine-tuning on a 12B decoder-only transformer (Gemma-3). The classification head variant adds a linear layer atop frozen embeddings; the instruction-tuning variant uses standard supervised fine-tuning on prompt-completion pairs. The compute envelope—single 24 GB GPU for both training and inference—is the key engineering constraint and the paper's practical selling point. No exotic hardware, no multi-node training. Integrity is mixed. The 2×2 factorial design with five training sizes is rigorous for an ablation study. The 100-report expert-adjudicated test set is well-constructed. But there's no pre-registration, no public benchmark beyond the authors' own annotated set, and no independent replication. The GPT-4o baseline is a moving target—model versions change, and the paper doesn't specify which snapshot was used. The statistical test (p=1.000 on the GPT-4o comparison) suggests a permutation or bootstrap test, but the small test set means confidence intervals are wide. The successor experiment the authors didn't run is scaling to other radiology tasks beyond ICH—chest X-ray findings, abdominal pathology, musculoskeletal reads. The 2×2 design could be replicated identically on five more extraction tasks in a month. The honest read: they're saving it for a follow-up paper that demonstrates generalizability, because a single-task result with 100 test cases is a proof of concept, not a deployment validation.