Imagine you're navigating a hilly landscape in dense fog. A first-order optimizer like Adam is a compass — it tells you which direction is downhill, but nothing about how steep or curved the terrain is. A second-order method is a topographic map: it knows the curvature, so it takes fewer, smarter steps. The problem is that in deep learning, the terrain has saddle points — places where the curvature goes negative — and your topographic map becomes a liar, telling you to go uphill. For decades, practitioners have patched this with ad hoc fixes: clipping negative eigenvalues, adding damping terms, running expensive line searches. SoftServe's core move is to derive curvature estimates that are positive-definite by construction, even when the true Hessian has negative eigenvalues, using a variational objective from Berglund et al. (2025). The committed claim: a family of quasi-Newton methods that scale to 136M-parameter networks and beat Adam, Muon, and SOAP on severely ill-conditioned problems — recurrent networks, deep autoencoders, physics-informed neural networks (PINNs), and a physics-informed diffusion model — without line searches or curvature corrections. This isn't a toy demonstration; 136M parameters puts SoftServe in the range of practical models. The architectural trick is twofold. First, SoftServe uses diagonal and Kronecker-factored approximations to the inverse Hessian, which is the same memory-efficiency playbook as K-FAC and SOAP but applied to a different curvature estimator. Second, all required matrix operations (inversions, square roots) use the coupled Newton-Schulz iteration — a fixed-point method that replaces eigendecompositions with sequences of matrix multiplications. This is the same trick Muon uses for its orthogonalization step. The result is an optimizer that's GPU-native: no custom LAPACK calls, no eigensolvers, just matmuls that tensor cores eat for breakfast. On the ladder: the paper benchmarks against Adam, Muon, and SOAP — all current-generation optimizers with active user bases. The claim is 'often achieving lower losses,' which is carefully hedged. The sweet spot appears to be ill-conditioned problems where first-order methods plateau. On standard well-conditioned tasks (think vanilla image classification), the advantage likely narrows or disappears. The paper does not claim universal superiority, which is the right move — the interesting question is whether the class of problems where second-order methods shine is large enough to matter in practice. Integrity considerations: the benchmarks include recurrent nets, deep autoencoders, PINNs, and a 136M-parameter diffusion model — a reasonable spread of notoriously ill-conditioned problems. But validation is same-team empirical. No independent replication exists yet. The choice of problems favors SoftServe's strengths (ill-conditioning), which is legitimate framing but worth noting. We don't know how SoftServe performs on the bread-and-butter tasks where Adam dominates (large-scale language modeling, standard vision). The absence of those benchmarks is the loudest silence in the paper. The milestone question for second-order deep learning optimizers has always been: can you match first-order wall-clock time at billion-parameter scale while delivering meaningfully better final loss? SoftServe reaches 136M. The next threshold is ~1B parameters on a standard training pipeline (language modeling or vision transformer). If the Kronecker-factored variant maintains its advantage at that scale without blowing up memory or per-step cost, the optimizer conversation changes. That's probably 1-2 years away given current hardware trends. The obvious experiment not run: training a large language model. Every optimizer paper eventually has to answer the LLM question because that's where the compute dollars are. The honest read is (a) compute budget — a 1B+ LLM training run is expensive, and this is a methods paper establishing the framework. But the authors surely know this is the question that will determine adoption. Expect it in a follow-up or from an industrial lab that picks up the method.