You've tuned every instruction in your inner loop. The assembly is hand-scheduled. The register allocator is perfect. You're done, right? This paper says you've been polishing the engine while ignoring the transmission. Imagine a Formula 1 team that obsesses over engine horsepower but never aerodynamic drag — at some point the marginal gains from the engine plateau and the car's shape becomes the bottleneck. That's exactly the argument here: post-quantum cryptography on embedded chips has focused almost entirely on the arithmetic kernels (the engine), while the execution environment — memory placement, cache hierarchy, peripheral DMA, clock configuration — remains largely unexamined (the drag). The committed claim: starting from a state-of-the-art SLOTHY-optimized ML-KEM implementation on Arm Cortex-M7, system-level optimizations alone yield up to 2.5% cycle reduction without any public-data tricks, and a selected public-data-reuse profile slashes encapsulation by 74.6% and decapsulation by 58.8%. These are not algorithmic changes. The wire format is untouched. The cryptographic guarantees are identical. The gains come entirely from how the surrounding system handles data movement and execution context. The baseline matters: SLOTHY is the current best-in-class instruction scheduler for Arm targets, producing assembly that is already aggressively optimized at the instruction level. By starting from SLOTHY output rather than a naive C compilation, the authors ensure their gains are genuinely additive — not substituting for work a better compiler would have done. They cover all three ML-KEM parameter sets (ML-KEM-512, -768, -1024), which prevents cherry-picking a single favorable configuration. Architecturally, the method is a layered profiling approach, not a single technique. The authors enumerate five system-level knobs: tightly coupled memory (TCM) placement to bypass cache latency, DTCM/ITCM configuration to keep hot data and code in single-cycle SRAM, peripheral DMA integration, clock-tree configuration for deterministic timing, and deterministic reuse of public encapsulation data across sessions. The big numbers come from the last knob — if you can cache a peer's public key and its derived data across multiple encapsulations, you skip the most expensive transforms entirely. This is a deployment constraint, not a cryptographic one: it's valid when the same public key is reused, which is common in TLS session resumption and IoT sensor networks. Integrity is solid but bounded. All measurements are cycle counts on real Cortex-M7 hardware (STM32 family), not simulation. The authors compare directly against SLOTHY-optimized baselines with explicit numbers for all three parameter sets. However, this is a single-team evaluation on a single MCU family. No independent replication exists yet, and the public-data-reuse profiles depend on deployment assumptions (key reuse frequency, threat model for side-channel leakage from caching public data) that are acknowledged but not deeply stress-tested. The milestone framing is practical and near-term. Post-quantum migration deadlines are approaching (NIST's 2035 deprecation target for classical algorithms in federal systems), and embedded devices — smart cards, IoT sensors, automotive ECUs — are the hardest deployment targets because they have the tightest cycle budgets. A 58-74% reduction in encapsulation/decapsulation cost on Cortex-M7 is the difference between 'ML-KEM fits in our power envelope' and 'we need a hardware upgrade.' The next concrete milestone is demonstrating these system-level gains on Cortex-M4 (the workhorse of constrained IoT) and on RISC-V targets, where the memory hierarchy is different. The obvious experiment the authors didn't run: side-channel analysis of the public-data-reuse profiles. Caching derived public-key material across sessions changes the timing and power profile of the device in ways that could leak information. The authors acknowledge this constraint but don't measure it. Most likely reason: side-channel evaluation requires specialized equipment and is a full paper on its own. This is a (c) — saving it for the next paper — and the FPS-2026 venue confirms the trajectory.