Imagine you're studying for an exam. You could re-read the same chapter ten times, or you could read four different chapters once each. If the test covers material from all four chapters, the second strategy obviously wins — yet most students (and most climate downscaling pipelines) default to the first. This paper measures exactly how much that matters for kilometer-scale weather prediction, and the answer is striking: what you train on matters overwhelmingly more than how much you train on. The committed claim: held-out error in a climate downscaling model grows linearly with climatological distance to the training data (RMSE = 0.83 + 2.95d), and this single relationship explains 90% of error variance — versus just 7% for training volume. That's not a tuning insight. That's a governing law. It means a group with modest compute can match the accuracy of one burning 4× more simulation budget, simply by choosing training months that span the target climate. The vehicle is CASPER, a U-Net with a structure-preserving loss that downscales 32 km ERA5 reanalysis to 1 km temperature, humidity, and wind fields. The architecture itself is not the contribution — U-Nets are standard workhorses in spatial downscaling. The contribution is the systematic experimental design: 24 configurations of one to eight training months, evaluated against held-out months and real station observations during documented heat waves. On extreme summer weeks, CASPER preserves fine-scale spatial structure and cross-variable physics that matched-budget baselines degrade. Against station data during heat waves, errors stay within 1.8 K. The integrity story is solid for a first demonstration but has clear boundaries. Validation uses held-out months and real station observations — not just internal simulation consistency. The 1.8 K station agreement during heat waves is a strong anchor. But the geographic transfer result is the honest caveat: moving to Vancouver without local data balloons error from ~1.3 K to 3.8 K. Eleven days of local simulation closes that gap, but 'how much local data for a new city?' is still an open empirical question, not a solved one. The milestone framing is where this gets practically important. Kilometer-scale urban heat modeling has been locked behind expensive dynamical downscaling simulations — the kind that require HPC clusters and weeks of wallclock time. If the error-distance scaling law holds across more regions and climate variables, it means small research groups and city planning offices can produce actionable heat-risk maps by strategically choosing a handful of simulation months rather than running exhaustive multi-year simulations. The 4× efficiency gain is not an asymptotic promise; it's demonstrated in the paper's own experimental grid. The obvious experiment not run: testing CASPER on a genuinely different climate zone (tropical, arid, high-altitude) rather than the Montreal-Vancouver Canadian corridor. The authors show geographic transfer degrades and that local data helps, but we don't know if the linear scaling law itself transfers — whether d predicts error with the same slope in, say, Phoenix or Lagos. The honest read is probably (a): compute and data access limited them to Canadian cities where they had WRF simulations ready. The scaling law's universality is the load-bearing question for the next paper. For practitioners in urban climate adaptation, the takeaway is immediate and actionable: audit your training data for climatological coverage before requesting more simulation. For the ML-for-climate community, the deeper takeaway is methodological — this paper demonstrates that data-efficiency scaling laws, analogous to those in large language models, exist in physical downscaling and can be characterized with modest experimental budgets.