Imagine you're baking bread, and instead of adjusting the oven temperature by trial and error — bake, check, rebake — you could trace the exact chain of cause and effect backwards from the final loaf to every decision you made: flour ratio, kneading time, oven temp. You'd know precisely which knob to turn and by how much. That's what automatic differentiation does for physics simulations, and Suêtes is the first system to make it work end-to-end for a full-featured regional weather model. The committed claim: Suêtes is a fully differentiable, non-hydrostatic, limited-area atmospheric dynamical core that supports reverse-mode automatic differentiation through every component — terrain geometry, time integration, boundary conditions, and physical parameterizations. This isn't a toy 2D demo. It includes both a semi-implicit semi-Lagrangian (SISL) scheme and a Split-Explicit Runge-Kutta (SERK) scheme with acoustic substepping, covering the range from coarse regional to convection-permitting scales. You get gradients of forecast outputs with respect to initial conditions, boundary fields, physical parameters, topography, and computational geometry — all in a single backward pass. Where does this sit on the ladder? The relevant predecessors are NeuralGCM (Kochkov et al., 2024) for differentiable global dynamics and various adjoint-based systems in operational NWP (like ECMWF's 4D-Var). NeuralGCM demonstrated differentiable global dynamics but used the hydrostatic approximation and spectral methods — fundamentally global-scale, not regional, not convection-permitting. Traditional adjoint models in operational weather (WRF, ICON) are hand-coded, brittle to maintain, and typically limited to specific components. Suêtes doesn't beat these on forecast accuracy — that's not the point. It beats them on flexibility: you get exact gradients through the entire model stack for free via JAX's autodiff, without hand-deriving or maintaining a separate adjoint code. The architecture sits in the differentiable physics simulation family, implemented in JAX for GPU/TPU execution. Two time-stepping schemes (SISL and SERK) give it range: SISL for larger time steps at coarser resolution, SERK for explicit convection-resolving dynamics. The terrain-following coordinate system, lateral boundary treatment (Davies relaxation with sponge layers), and included parameterizations (Kessler microphysics, Smagorinsky diffusion) are all written to be AD-compatible. This is a deliberate engineering choice — every branching, every interpolation, every implicit solve was written or rewritten to avoid AD-hostile patterns. The validation regime is a suite of six numerical experiments, not a single benchmark number. These include a tracer-source inverse problem (can gradients recover a hidden emission source?), sensitivity analysis of a 3D squall line, adversarial perturbation construction for an ERA5-driven downslope hurricane-force wind event (the suêtes winds of Cape Breton — hence the name), topography optimization, and terrain-following coordinate optimization. Each experiment demonstrates a different use of the gradient information. The experiments are self-consistent demonstrations, not comparisons against an external benchmark — this is appropriate given that no directly comparable differentiable limited-area non-hydrostatic core exists. The milestone question is about scale and operational relevance. The demonstrated experiments are research-scale proofs of concept. The next concrete threshold is coupling Suêtes with a learned physics parameterization (neural network replacing Kessler microphysics, say) and training it end-to-end on reanalysis data at convection-permitting resolution (~1-4 km) over a domain large enough to matter for regional forecasting — think a full province or small country for a multi-day forecast. That's where differentiable regional dynamics becomes a competitor to operational limited-area models, not just a research tool. The obvious experiment they didn't run: end-to-end training of a hybrid model with neural parameterizations on real observational or reanalysis data, evaluated against operational NWP baselines. The likely reason is (a) — compute and data pipeline cost. Building the differentiable core is the hard engineering; coupling it with learned components and scaling it is the next paper. The framework is explicitly designed for this use case, and the authors signal it clearly. This is a platform paper, not a results paper, and it's honest about that.