You know that friend who insists on printing out Google Maps directions even though you both have phones with GPS? Abstract Meaning Representation (AMR) for large language models is turning out to be that friend. AMR — a structured graph that maps the 'who did what to whom' of a sentence — was genuinely useful when NLP systems couldn't figure out relational structure on their own. The question this paper answers is whether modern LLMs still need that crutch. The answer is no, and the path to getting there is more interesting than the destination. Nguyen, Staiano, and Sullivan attempted to reproduce recent work claiming substantial downstream gains from feeding AMR graphs alongside text into LLMs. They couldn't. When they standardized hyperparameter selection — using a consistent, unified protocol instead of the original papers' bespoke tuning — text-only baselines consistently matched or exceeded AMR-augmented models. The prior gains were artifacts of experimental setup, not genuine signal from the structural representation. This is a clean negative result, and the authors don't stop at 'it didn't work.' They introduce a perplexity-based probe that directly measures whether AMR provides relational knowledge an LLM doesn't already have. The probe asks: when you show the model AMR-encoded relational content, does the model's uncertainty about the sentence decrease? It doesn't. The LLM already knows who did what to whom. AMR is telling it things it already knows. The architectural logic is straightforward. Early NLP systems (pre-BERT era and even early fine-tuned models) had limited capacity to extract relational structure from raw text. AMR provided an explicit scaffold. But modern LLMs, trained on billions of tokens with attention mechanisms that naturally learn dependency structures, have internalized exactly the kind of relational reasoning AMR encodes. The augmentation is redundant — not wrong, just unnecessary. The integrity story here is strong for a reproduction study. The authors use 18 tables and 6 figures across 23 pages — this is exhaustive, not cherry-picked. The key methodological contribution is the unified hyperparameter protocol: they show that when you let each model variant pick its own best hyperparameters (as prior work effectively did), you can make AMR augmentation look beneficial simply because it got luckier in the tuning lottery. Control for that, and the signal vanishes. The broader implication lands squarely in an active field fight: should we keep layering structured linguistic representations onto LLMs, or have scale and pretraining made explicit linguistic structure obsolete? This paper is ammunition for the 'scale subsumes structure' camp. It doesn't prove that NO structured representation could ever help an LLM — only that AMR, the most mature and widely-studied semantic representation in NLP, provides nothing the model doesn't already have. What makes this paper valuable isn't the novelty of the claim — many practitioners already suspected this — but the rigor of the demonstration and the diagnostic probe. The perplexity-based measurement gives the field a reusable tool for asking 'does this augmentation actually add information?' before running expensive downstream experiments. That's the lasting contribution: not the negative result itself, but the methodology for detecting redundant augmentation.