Imagine you've been building a cathedral for decades — hand-cutting every stone with Navier-Stokes equations, running massive parallel simulations on supercomputers to predict whether it will rain on Zurich tomorrow. Now someone shows up with a 3D printer that produces a cathedral of roughly the same quality in a fraction of the time, trained by studying photographs of thousands of existing cathedrals. That's what MeteoSwiss just did with Varda-single-1.0: replaced the physics-based stone-cutting with a learned model that ingests 20 years of high-resolution reanalysis data and outputs hourly 1 km forecasts across Switzerland's brutally complex Alpine topography. The committed claim: a data-driven weather model can produce deterministic regional forecasts at 1 km resolution over mountainous terrain that are competitive with — and in many cases better than — the operational numerical weather prediction (NWP) systems it might one day replace. This is not global-scale ML weather (Pangu-Weather, GraphCast, GenCast operate at 25-31 km). This is kilometre-scale, regional, over mountains — where NWP has traditionally held its strongest advantage because fine-grained terrain forcing matters enormously. The architecture is a stretched-grid Graph Transformer in the Anemoi framework, using an encoder-processor-decoder structure. Two models work in tandem: a 6-hourly autoregressive forecaster and a temporal downscaler that fills in hourly steps between the forecaster's jumps. The training curriculum is a three-stage ladder — pre-training on ERA5 global reanalysis (31 km), then training on a 20-year km-scale regional reanalysis (KENDA-1), then fine-tuning on operational km-scale analyses. This curriculum is the structural bet: it assumes that coarse global weather patterns transfer usefully to fine Alpine dynamics, which is not obvious over terrain where valley winds and convective cells are dominated by local orography. Against the ladder, the results are genuinely strong. Varda-single broadly matches ICON-CH1-EPS (MeteoSwiss's 1 km physics-based control run) out to +33 hours and generally outperforms ICON-CH2-EPS (the 2 km control) out to +120 hours. For most headline scores and variables — temperature, pressure, humidity — the ML model is competitive. But the honest caveats matter: local wind maxima are underestimated, and convective precipitation fields come out too smooth. Both failures are the signature of squared-error (MSE) training, which penalizes being wrong more than it rewards being sharp. The case studies investigating local wind representation over complex terrain reveal this is not just an aggregate stat — it's a structural weakness for safety-critical applications like wind warnings. The integrity picture is solid for an ML weather paper. Verification runs over a full year against both operational analyses and surface station observations. The baselines are real operational systems, not stale or cherry-picked. Model weights are released on HuggingFace. The authors are explicit about where the model fails. What's missing: no pre-registration (standard for this field), and no independent replication yet. The smoothing bias from MSE training is acknowledged but not addressed — this is the obvious next experiment (probabilistic or generative training) that was not run. The milestone question is about resolution and reliability, not just accuracy. At 1 km, the model enters the domain where individual thunderstorms and valley wind systems should be resolvable. The next concrete threshold is whether a probabilistic or ensemble version of this architecture can match the spread calibration of ICON-CH1-EPS ensembles — not just the deterministic control. If it can, the compute savings could be transformative: an ML ensemble that costs orders of magnitude less than a physics-based ensemble run on a supercomputer. The successor experiment the authors didn't run is generative or diffusion-based training to address the smoothing problem. The paper acknowledges the MSE limitation explicitly. The honest read: this is almost certainly being saved for the next paper (Varda-ensemble or Varda-2.0). A secondary gap is testing on truly extreme events — the case studies probe weaknesses but the verification period is one year, which may not contain the tail events that matter most for civil protection.