Imagine you're a detective with two filing cabinets. One holds everything about the suspect — photos, phone records, financial statements, DNA. The other holds encyclopedic knowledge about how crimes typically work — motives, methods, forensic science. Each cabinet alone gives you hunches. But the cases you actually crack are the ones where you draw a literal line on a corkboard from a specific piece of suspect evidence to a specific fact in the encyclopedia. MM-KG is that corkboard, and the paper's core finding is that the lines themselves — not the cabinets — are what matter. The committed claim: clinical LLMs benefit from knowledge graphs not as passive background context (the way most KG-augmented systems use them) but specifically and only when explicit typed edges link a patient's multimodal observations to the biomedical relation a question requires. The authors back this with a clean 2×2 factorial design — patient evidence alone, knowledge alone, both, neither — and show that on questions demanding both sources, neither alone performs meaningfully above chance, while their interaction yields AUROC gains of +0.194 on MIMIC-IV and +0.299 on ADNI. That interaction term is the paper's load-bearing number. The architecture is a layered typed graph. Modality-specific 'harmonizers' convert EHR text, imaging, genomic, and biospecimen data into typed observation nodes mapped to UMLS concepts. A route-prioritized aligner then links these to a biomedical knowledge graph (think: drug-treats-disease, gene-associated-with-phenotype edges). At query time, a conditioned retrieval step selects a compact subgraph — not the whole KG — for an LLM or GNN to reason over. This last piece is crucial: query-conditioned retrieval hits 0.731 AUROC with 6.8× less context than the strongest generic retrieval policy. Static KG dumps, by contrast, show no consistent gain on outcome prediction. The method is saying: targeted retrieval over explicit links beats brute-force context stuffing. On the ladder, MM-KG outperforms MindMap by +0.131 Hits@1 on five-candidate ranking and leads an adapted GraphCare on items that require consulting the patient record. The ablation is particularly convincing: deleting the single answer-bearing relation from the retrieved subgraph returns Hits@1 to the no-knowledge baseline. This is rare in KG-augmented work — most papers show the KG helps but cannot show which specific edge was load-bearing. The causal traceability here is a genuine methodological contribution. Integrity is mixed. The 2×2 factorial design is strong methodology — it separates patient evidence, biomedical knowledge, and their interaction cleanly, which most competitors don't attempt. The use of MIMIC-IV and ADNI are community-standard datasets. But there's no pre-registration, no independent replication, and the benchmarks (five-candidate ranking, drug-controlled AUROC) are somewhat bespoke. The ablation where deleting the answer-bearing edge collapses performance is the strongest internal check, but it's self-graded — the authors chose which edge to delete. Code availability is not mentioned. The milestone question is about scaling this to real clinical deployment. Two datasets with curated UMLS mappings is a proof of concept. The next concrete threshold is demonstrating MM-KG on a full hospital EHR system with hundreds of thousands of patients and real-time query latency — say, sub-second retrieval over 500K+ patient graphs. The gap between 'works on MIMIC-IV' and 'works in production' is where most clinical AI dies. The harmonizers' ability to handle messy, incomplete, contradictory real-world records is entirely untested. The obvious experiment not run: applying MM-KG to a prospective clinical decision support task where a physician actually uses the retrieved subgraph and the system's traceable edges to change or confirm a decision. The authors built traceability into the architecture — you can point to exactly which edge mattered — but never tested whether that traceability is useful to a human clinician. My read: this is being saved for the next paper, possibly with a clinical collaborator. The current work is the infrastructure paper; the clinical validation paper is the sequel.