Imagine you're redecorating a house and you've decided you can only mix paint from three specific cans. You spend hours finding the perfect ratio of Can A, Can B, and Can C. Then someone points out: you could just go to the paint store. The optimal color was never a blend of those three cans — it was always available, you just never looked outside the constraint you imposed on yourself. That's this paper's core move, applied to model merging. The committed claim: the standard practice of searching for linear-combination coefficients over task-specific weight updates imposes an implicit regularization — restricting the solution to a low-dimensional subspace — and this regularization actively hurts multi-task performance. Removing it and optimizing directly in full weight space consistently improves results across architectures (ViT, CLIP, RoBERTa), domains (vision, NLP), and even in extreme data-poverty (one sample per class). This is not a new merging algorithm; it's an argument that the entire pipeline's search space is wrong. The ladder here is interesting. The paper doesn't claim to beat every merging method on every benchmark — it claims something more structural. It shows that Task Arithmetic, Ties-Merging, and DARE all improve when you remove the subspace constraint and fine-tune in full weight space. More embarrassingly for the field, directly fine-tuning the pretrained model (no merging at all) outperforms some existing merging methods. The baselines are current and named: Task Arithmetic (Ilharco et al. 2023), Ties-Merging (Yadav et al. 2024), DARE (Yu et al. 2024). The authors are honest that some merging methods still win when the additional dataset is used optimally — but the gap narrows dramatically. Architecturally, this is a meta-study of the optimization landscape, not a new model family. The key insight is geometric: task-specific weight updates (τ vectors) span a subspace of dimension equal to the number of tasks — often 7 or 8 dimensions in a space with millions. Searching only over coefficients λ₁...λₖ means you're optimizing over a hyperplane in a space that has vastly better solutions elsewhere. The paper uses standard SGD/Adam to explore the full weight space, starting from the merged initialization. The compute cost is modest — it's fine-tuning, not pretraining. Integrity is solid but not exceptional. The experiments span eight vision tasks (SUN397, Cars, DTD, EuroSAT, GTSRB, MNIST, RESISC45, SVHN), multiple NLP benchmarks, multiple architectures, and the extreme one-shot regime. The comparison baselines are current. However, the evaluation is entirely self-conducted — no independent replication, no pre-registration. The benchmarks are standard community benchmarks, which helps. The key risk: the paper's core result (unconstrained fine-tuning beats constrained coefficient search) is almost tautological given enough data and compute. The paper's actual contribution is showing how little data you need and quantifying the gap. The milestone question is less about a single number and more about a paradigm shift. If this result holds up, the next milestone is: at what model scale does the subspace constraint become beneficial again? There may be a crossover point where the regularization helps — when the additional dataset is truly tiny relative to the weight-space dimensionality. The paper partially addresses this with the one-shot experiments but doesn't map the full frontier. The obvious experiment not run: scaling to LLMs with billions of parameters, where the gap between subspace dimensionality and full weight space is even more extreme but compute for full fine-tuning becomes prohibitive. The honest read: this is (a) compute-limited. Fine-tuning a 70B model across eight tasks requires serious hardware. The authors likely know this is the next paper — it's the obvious extension that would make the result impossible to ignore.