Imagine you're a burglar who knows a homeowner's security cameras have a slow drift — the viewing angle shifts a few degrees every hour between recalibrations. A static burglar picks one blind spot and hopes. A smart burglar watches the drift and adjusts her path each hour. This paper asks: how much does the smart burglar actually gain in quantum key distribution? The committed claim is that adaptive eavesdropping on drifting QKD channels can be formalized as a constrained Markov decision process, and that a reinforcement-learning agent operating in this framework reaches 98–99% of the provable upper bound on extractable information. The attacker doesn't just optimize rotation angles on a fixed gate template (as Decker et al. did) — it jointly searches gate structure and parameters, producing compact circuits that form a discrete action set. This is the genuine advance: extending learnt attacks to noise models where no analytical template exists, including the amplitude damping channel. The ladder here is unusually clean. On device-independent E91 under bilateral depolarizing noise, the best fixed (non-adaptive) circuit achieves Holevo information of 0.135. The RL attacker reaches 0.348 at zero detection — a 2.6× improvement — capturing 98% of the dynamic-programming upper bound. On BB84 under a drifting bit-flip channel, the attacker exceeds a conservative noise-indexed rule by 0.024 in fidelity, hitting 99% of the upper bound. Critically, when started from random gate sequences (no analytical priors), the search recovers the known optimal cloners and matches the collective-attack key rate. This self-consistency check is the strongest validation the paper offers. Architecturally, this is a constrained MDP with an Ornstein–Uhlenbeck noise process driving the channel state and an abort condition functioning as a budget constraint over each block of rounds. The action space is a discrete set of compact quantum circuits found by joint topology-parameter search — a sparse search method that earned a parallel NeurIPS 2026 acceptance. The RL agent selects one circuit per round. The value of adaptation is sandwiched between the best fixed circuit (lower bound) and a dynamic-programming upper bound, which gives the analysis unusual rigor for an ML-meets-quantum paper. Integrity is solid for a theoretical/simulation paper. The authors bound their attacker's performance from above and below, recover known analytical results without seeding them, and explicitly compare against both fixed-circuit and noise-indexed baselines. The drifting-channel model (Ornstein–Uhlenbeck) is physically motivated, not cherry-picked. The main limitation is that all validation is simulation-based — no physical QKD hardware is involved. The per-basis vs. averaged error-rate constraint analysis showing the gain from basis asymmetry changes sign is a nice honesty move; many papers would quietly drop the less favorable comparison. The milestone question is where this gets practical. Current QKD security proofs assume stationary channels and provision key rates accordingly. If adaptive eavesdroppers can extract 2.6× more information by following device drift, the gap between assumed security and actual security under realistic deployment is non-trivial. The next concrete step is demonstrating this on a physical QKD testbed with real device drift, which would force the security-proof community to update finite-key analyses for non-stationary channels. The authors note QCrypt 2026 presentation, signaling the community is paying attention. The obvious unrun experiment is the physical one: deploying this adaptive attack framework against a real QKD link with measurable device drift. The honest read is (a) — they lack access to or budget for a suitable QKD testbed. The parallel NeurIPS paper on the sparsity/search method suggests the team's core expertise is ML, not experimental quantum optics. A second gap: scaling the MDP framework to decoy-state or continuous-variable QKD protocols, which dominate deployed systems. That's likely (c) — the next paper.