Imagine you're a home cook trying to figure out whether your steak is better because of the pan, the heat, or the seasoning. You'd never change all three at once — you'd swap one variable at a time. That's the core mechanism here: the authors built a modular test kitchen for quantum circuit optimizers, separating the quantum part (how you estimate the search direction) from the classical part (how you update the parameters), so you can finally see which piece matters for which dish. The committed claim is modest but genuine: no single optimizer dominates across all quantum workloads, and the interactions between quantum estimation strategy and classical update rule are workload-dependent in ways that existing papers conflate. This is a benchmarking infrastructure paper, not a breakthrough algorithm paper. The authors are explicit about this — they say their hardware runs 'do not establish an optimizer ranking or isolate a causal effect of device noise.' That level of honesty is rare and welcome. The framework covers four workloads: QAOA for MaxCut combinatorial optimization, VQE for molecular hydrogen ground-state energy, Iris flower classification (a standard ML benchmark), and QCNN for binary MNIST digit classification. Optimizers tested span gradient-based (SPSA, parameter-shift), stochastic, and derivative-free families. Evaluation happens under finite-shot simulation and on a 156-qubit IBM processor. The paper separates terminal objective values from best-observed values — a distinction most benchmarks ignore but which matters enormously when your evaluation budget is limited. The architecture contribution is the modular decomposition itself. By factoring the optimization loop into a quantum search-direction oracle and a classical parameter-update rule, the framework makes it possible to mix and match components. This is the right abstraction for a field where hardware noise, shot budgets, and parameter count all interact in non-obvious ways. The paper belongs to the variational quantum algorithm family (VQAs), sitting alongside QAOA, VQE, and variational QML as applications. Integrity is mixed. On the positive side, the authors run multiple seeds on simulator benchmarks (three seeds, reporting both peak and mean performance), test across genuinely different workloads, and are candid that hardware results are illustrative rather than conclusive. On the negative side, there's no pre-registration, no independent replication, and the hardware runs are limited case studies rather than systematic sweeps. The paper doesn't cherry-pick a winner, but it also can't establish one — which is more a function of the noisy hardware than the methodology. The milestone question for this line of work is clear: when can we run systematic optimizer benchmarks on hardware with enough qubits and low enough error rates that the optimizer ranking is stable and meaningful? The paper runs on 156 qubits but can't draw causal conclusions about device noise effects. That's the gap. Error rates below ~0.1% per gate on 200+ qubits would likely make hardware benchmarking decisive rather than illustrative. The obvious next experiment not run: scaling the parameter count within a single workload type to find where optimizer rankings flip. The VQE (H₂) is tiny; the QAOA instances are modest. Testing at 50-100 parameters with systematic noise injection would stress-test the framework's core promise that interactions are workload-dependent. Best guess on why it wasn't done: compute budget. Running a full factorial design across optimizers × workloads × parameter scales × shot budgets on real hardware is expensive, and this paper already covers a lot of ground for what reads like an academic-group effort.