Imagine you're packing for a long trip with a strict weight limit. The usual strategy: go room by room, tossing whatever looks least important from each drawer. That works when you're cutting 10% — but at 50% or 70%, the room-by-room approach starts throwing out things that matter deeply to OTHER rooms. Your kitchen knife goes because the kitchen drawer was heavy, but now you can't prepare anything you packed from the pantry. The smarter move is to lay everything on the bed, consider how items depend on each other across the whole trip, and cut globally. That's what Learnable Subspace Projections (LSP) does for neural network weight compression. The committed claim: existing low-rank compression methods fail at high compression ratios because they choose what to discard using local, per-layer criteria — activation energy, reconstruction error, quadratic loss approximations — that ignore how errors compound through the network's depth. LSP replaces these local heuristics with a single global optimization: learn an orthogonal projector for each layer (or tied group of layers), optimize all projectors jointly against KL divergence to the dense model's output distribution, and keep the original pretrained weights frozen throughout. After training, the projectors fold into standard low-rank matrix factors with zero inference overhead. The architecture is clean and specific. Each weight matrix gets decomposed via whitened SVD truncation as initialization, then an orthogonal projector is optimized end-to-end. Rank allocation across layers is guided by the output KL each projector induces per parameter saved — a principled budget-aware criterion rather than uniform rank assignment. A key structural trick: in attention layers, tied groups of layers that read the same activations share one projector factor, which means key-value caches can store one narrow latent instead of full keys and values. This is where the 13.5× combined weight + KV cache compression at 128k context comes from, versus at most 6.5× for untied baselines. The ladder position is strong and honestly reported. At 70% compression on Llama-2-7B, LSP achieves 10.9 WikiText-2 perplexity and 42.2% mean zero-shot accuracy. The strongest baseline (presumably ASVD or SliceGPT-family methods, given the local-criterion description) manages only 13.3 perplexity and 36.0% accuracy — a 2.4 perplexity point gap and 6.2 percentage point accuracy gap. Critically, the paper claims this advantage WIDENS as compression increases. The method is also tested across model scales (OPT-125M, OPT-1.3B, Qwen3-4B, Llama-2-7B) and modalities (ViT-B/16), which guards against cherry-picking a single favorable configuration. Integrity is solid but not airtight. WikiText-2 perplexity is a community-standard benchmark, and zero-shot accuracy across multiple tasks is the accepted evaluation for compressed LLMs. The models tested span three families and two modalities, reducing the risk of method-architecture coupling. However, the paper does not appear to be pre-registered, no independent replication exists yet, and the baseline naming is vague in the abstract — we're told 'strongest baseline' without the explicit method name, which always deserves a raised eyebrow. The 1.6× decode speedup claim is limited to small batch sizes, an honest qualifier that prevents overclaiming. The milestone picture is concrete and practical. Today's result: 70% compression with ~2 perplexity points of damage on a 7B model, plus 1.6× decode speedup and 13.5× cache compression at 128k context. The obvious next threshold is 70-80% compression on 70B+ models (Llama-3-70B, Qwen-72B) while maintaining single-digit perplexity damage — which would make this method directly deployable for edge inference and long-context serving at scale. The gap is probably 1-2 years given compute scaling and architecture refinements. The experiment the authors did NOT run is the most revealing absence: scaling to 70B+ parameter models and testing on harder compression targets (80-90%). The honest read is option (a) — compute budget. Training orthogonal projectors end-to-end against a global KL objective at 70B scale is expensive, and the paper demonstrates the principle convincingly at 7B. Expect the 70B experiments in a follow-up within 12 months. A second missing experiment: combining LSP with quantization (e.g., GPTQ or AWQ), which would be the practical deployment configuration. This is likely being saved for the next paper.