Imagine you're assembling IKEA furniture. You have two strategies: follow the printed instructions step by step, or just eyeball it using general carpentry instincts — keep things level, tighten screws evenly, don't force joints. For a simple bookshelf, both approaches get you to the same place. But for one of those baroque corner desks with angled shelves and hidden cable channels, the specific instructions start pulling ahead of vibes-based carpentry. Now imagine someone hands you instructions for a desk that's been redesigned since printing — the holes don't line up anymore. Following those stale instructions is now worse than winging it. That's this paper in a nutshell. The committed claim: incorporating a known governing-equation residual as a loss term consistently outperforms the best generic regularizer (weight decay, dropout, spectral normalization) for nonlinear PDE surrogates, but this advantage vanishes for linear PDEs and turns actively harmful when the computational grid under-resolves the physics. This is a separation result, not a new architecture. The paper's value is diagnostic — it tells you when to bother with physics-informed losses and when you're wasting effort or injecting bias. The experimental protocol is unusually disciplined for this subfield. Farazpay and Bora evaluate against a from-scratch FNO baseline under matched hyperparameter tuning budgets across five PDEs: heat, advection-diffusion (linear), and Burgers, KdV, Allen-Cahn (nonlinear). Fixed-capacity comparisons show the residual loss wins everywhere, but the capacity sweep is where the real story lives. As you scale model size, the residual's advantage grows for the three nonlinear equations but collapses to parity or below for the two linear ones. The linear PDEs are the bookshelf — generic regularization suffices. Nonlinear dynamics are the corner desk — you need the actual instructions. Two auxiliary results sharpen the picture. First, a pre-registered hypothesis that the residual's benefit would be activated only under data-sparse regimes is cleanly falsified: the residual helps even under full supervision for nonlinear problems. This is important because it means the mechanism isn't just 'the residual fills in where data is missing' — it's providing genuinely non-redundant structural information. Second, when the spatial grid is too coarse for the nonlinear dynamics, the naive PDE residual becomes a wrong constraint and actively degrades performance. The instructions-don't-match-the-holes failure mode. The paper also tests two fashionable priors — cross-family pretraining (train on one PDE family, transfer to another) and in-context conditioning — and finds neither outperforms the strong from-scratch baseline in the regime studied. These are not killed definitively; the regime is specific and the baselines are strong. But the result is a useful corrective to papers that claim transfer-learning benefits without controlling for a properly tuned single-task baseline. Methodologically, the use of pre-registered hypotheses (at least one is explicitly called out) and matched tuning budgets is above the norm for this corner of ML. The study is simulation-only, comparing neural surrogates against each other rather than against ground-truth experimental data, but this is the right validation regime for the question being asked. The question isn't 'can we solve PDEs' — it's 'does the physics prior carry non-redundant information inside the neural surrogate pipeline.' The practical upshot is a decision rule for practitioners building neural PDE surrogates: if your target PDE is nonlinear and your grid resolves the dynamics, pay the implementation cost of a residual loss — it's not just regularization, it's information. If your PDE is linear, save yourself the trouble. And if you're on a coarse grid, the residual will lie to your model.