Imagine you're a weather forecaster, but instead of predicting rain, your job is to predict where and how intensely the planet will be on fire next week. Right now, NOAA's aerosol forecast system does something almost comically crude: it grabs the most recent satellite snapshot of global fire activity — which is already 1.5 days stale by the time it arrives — and then holds that snapshot frozen for the next 5-7 days of forecast. It's like navigating rush-hour traffic with a photo of the highway taken yesterday morning. This paper builds the replacement: a spatiotemporal graph neural network that ingests recent fire observations, reanalysis meteorology, land cover, and vegetation data to predict fire radiative power (FRP) globally at 0.1° resolution, one to seven days ahead. The committed claim: a data-driven model can produce useful multi-day global FRP forecasts that substantially outperform persistence (the frozen-snapshot baseline). Trained on 2020-2022 GBBEPx data and evaluated on held-out 2023-2024, the model reduces mean squared error by 32% at day-1 and 43% at day-7 in 2023, and by 24% and 40% in 2024. At coarser 1° resolution, the critical success index (CSI) ranges from 0.32 to 0.60. The seasonal cycle is reproduced well. The model detects large fires reliably but systematically underestimates their intensity — a bias pattern the authors flag honestly. The architecture is a spatiotemporal graph neural network operating on a global grid. Each grid cell at 0.1° is a node; edges encode spatial adjacency. Temporal context comes from recent fire history plus reanalysis fields (wind, humidity, temperature — the usual suspects that govern fire spread). The target variable, GBBEPx FRP, is a NOAA operational satellite product derived from VIIRS and MODIS sensors. This places the method squarely in the geoscience-ML family: physics-informed features fed to a learned graph propagation model. The compute requirement is not discussed in detail, but the global 0.1° grid is roughly 6.5 million land cells, so this is a non-trivial spatial problem. The ladder question is where we need to be careful. The headline baseline is persistence — literally doing nothing and assuming fires don't change. This is operationally relevant because it's what NOAA's system effectively does, but it's also a weak baseline by ML standards. The paper does not compare against process-based fire spread models (e.g., FARSITE, Prometheus), nor against simpler ML baselines (random forests, linear models on the same features), nor against other recent deep-learning fire prediction efforts. The 32-43% MSE improvement over persistence sounds strong, but without a tougher competitor, we don't know how much of that is architecture versus simply having the right input features. Integrity has genuine strengths and notable gaps. The temporal split — train on 2020-2022, evaluate on 2023-2024 — is the right approach and avoids data leakage. Reporting separate results for 2023 and 2024 shows honest year-to-year variance (the 2024 numbers are weaker, as you'd expect from a model encountering new fire regimes). The authors are forthright about systematic intensity underestimation and small-fire placement errors. However, there's no code release mentioned, no ablation study isolating which input features drive skill, and no comparison to other ML methods. The paper is submitted to AI for Earth Systems, a relatively new journal — not yet a high-bar community benchmark venue. The milestone here is operational integration. The authors explicitly frame this as a stepping stone toward replacing frozen fire inputs in NOAA's GEFS-Aerosols system. The gap isn't accuracy per se — it's intensity calibration. The model detects fires well but undershoots their power, and aerosol forecasts need intensity, not just location. The concrete next number: close the intensity bias so predicted FRP can drive downstream PM2.5 forecasts with skill comparable to using actual observed FRP. That's probably 1-2 iterations away if the team has operational access. The obvious experiment not run is an ablation study showing how much skill comes from the GNN structure versus the meteorological inputs versus fire history alone. A simple gradient-boosted tree on the same features would tell us immediately whether the graph architecture is carrying its weight or whether 80% of the skill is in the features. My read: the team is focused on operational viability and building the full pipeline, not on ablation-level rigor — this is an applications paper, not a methods paper. That's a legitimate choice, but it means we can't separate architecture novelty from feature engineering.