Imagine you're teaching someone to play chess by showing them thousands of finished games — they learn patterns, but they don't learn the rules. Hand them a board position they've never seen and they freeze. Now imagine you also teach them how pieces move, that bishops stay on diagonals, that pawns can't go backward. Suddenly they can reason about novel positions. That's the core mechanism here: instead of training neural networks purely on quantum chemistry answers, the authors bake the governing equations of Kohn-Sham density functional theory directly into the model's learning objectives and architecture, so the network internalizes the rules of quantum mechanics, not just its outputs. The committed claim is sharp: by aligning machine-learned electronic ground-state models (GSMs) with the physics of KS-DFT — enforcing orthonormality of orbitals, removing optimization pressure on gauge-invariant degrees of freedom, and training against the Kohn-Sham equations themselves — you get models that generalize dramatically better out-of-distribution. The headline numbers are striking: on density-based GSMs extrapolating from QM9 (small molecules) to QM40 (larger ones), the combined ON-Loss and GROOT methods cut energy MAE by 79.1% and force MAE by 83.4% versus previous state-of-the-art. For Hamiltonian-based GSMs, ON-Loss plus ROCKET reduce energy and force MAEs by 99.8% and 95.9% respectively over the strongest baseline. The architecture family is equivariant graph neural networks predicting electronic structure quantities — either electron density or Hamiltonian matrices — that sit on a cost-accuracy Pareto frontier between dirt-cheap MLIPs and full Kohn-Sham DFT self-consistent field calculations. The key structural choices are three training innovations: ON-Loss enforces orthonormality of predicted molecular orbitals (removing the model's freedom to waste capacity on unphysical states); GROOT restricts density training to the occupied orbital manifold using Grassmann geometry; and ROCKET conditions Hamiltonian training on the Kohn-Sham eigenvalue equation itself while handling gauge freedom through optimal conditioning. These aren't just regularization tricks — they reshape what the model is asked to learn. The validation regime is solid but not bulletproof. The extrapolation experiments use a genuine out-of-distribution setup: train on QM9 (≤9 heavy atoms), test on QM40 (up to 40 heavy atoms). That's a real size-extrapolation challenge, not an interpolation dressed up as generalization. They also demonstrate on QMugs (drug-like molecules) with a self-consistency rejection filter that flags bad predictions — rejecting fewer than 0.4% while hitting 0.07 mHa energy MAE. Transfer to Transition1x (reactive chemistry, transition states) reaches sub-chemical-accuracy energy errors. The baselines compared include PhiSNet, QHNet, and recent density GSMs — these are current, not stale. The missing piece: no independent replication, no pre-registration, and code availability is not explicitly stated in the abstract. The milestone that matters is practical adoption in molecular dynamics and drug discovery pipelines. Today's MLIPs are fast but brittle outside their training domain; full DFT is reliable but expensive. GSMs occupy the middle ground — slower than MLIPs, faster than DFT, but now with dramatically better generalization. The next concrete threshold is whether these models can reliably replace DFT for production-scale conformer screening of drug candidates (thousands of molecules, diverse chemistries) without manual quality checks. The 0.07 mHa error on QMugs with <0.4% rejection rate is close but not there yet — you'd want <0.1% rejection across more diverse chemical spaces. The experiment the authors didn't run — and the one that would change the conversation — is scaling to periodic systems (crystals, surfaces, catalysis). All demonstrations here are on molecular systems. Periodic boundary conditions introduce different symmetry constraints and the density/Hamiltonian representations change fundamentally. The honest read: this is almost certainly (c), being saved for the next paper, since the mathematical framework (Grassmann manifold training, KS-equation alignment) generalizes in principle but requires substantial engineering for periodic codes. A second gap: direct comparison of wall-clock cost against actual DFT calculations, not just accuracy comparisons against DFT labels. The efficiency claim — 'computational costs between MLIPs and KS-DFT' — deserves concrete timing numbers. What makes this paper worth your time isn't any single number but the design philosophy. Most ML-for-science papers treat physics as a data source; this one treats physics as an inductive bias. The distinction matters: data-fitted models memorize the answer key, physics-aligned models learn the subject. As molecular ML moves toward real deployment in pharma and materials, the question of whether you can trust a model outside its training set is existential, and this paper offers a principled answer that actually moves the needle.