Imagine you're moving apartments and you pack every room into its own box — but then realize that the kitchen box and the dining room box are 90% identical plates. You wouldn't ship both. You'd merge them. VFold does exactly this for the value cache in large language models: it notices that adjacent transformer layers store nearly identical value states and folds them together, shipping one box instead of two. The committed claim: inter-layer value caches in LLMs exhibit a symmetry structure that can be exploited to merge caches across layers, cutting value cache memory roughly in half without modifying the model architecture or retraining. This is a training-free, inference-time compression that works as a drop-in on existing models. The key insight is that value representations across consecutive layers are far more similar than key representations — a structural asymmetry the field has underexploited. On the ladder, VFold positions itself against existing KV cache compression methods. Cross-layer sharing approaches like CLA and YOCO require architectural changes and retraining. Quantization methods like KIVI compress individual caches but hit walls at extreme ratios. VFold's value is that it's orthogonal to both families: you can merge value caches across layers AND quantize the remaining cache AND prune keys, composing compression ratios that no single method achieves alone. The paper demonstrates this on standard LLM benchmarks, showing that merging up to ~50% of value cache layers preserves task performance while stacking with 2-bit quantization. Architecturally, this lives in the attention-based KV cache management family — specifically the inference-time, post-hoc compression branch that doesn't touch training or model weights. The method relies on a structural property of transformers: value vectors in adjacent layers occupy similar subspaces, likely because residual connections dominate the update signal at deeper layers. The compute property being leveraged is that cosine similarity between consecutive value caches is high enough that merging introduces negligible error, while key caches don't share this property — hence the asymmetric strategy. Integrity is reasonable but bounded. The evaluation uses standard LLM benchmarks (likely perplexity and downstream tasks), and the baselines include current compression methods. However, this is self-graded — the authors propose the method and evaluate it. No independent replication exists yet, and the specific benchmark suite selection could favor the method. The composability claim (stacking with quantization and pruning) is the most interesting but also the most vulnerable to cherry-picking across compression ratio × task × model combinations. The milestone here is practical: current KV cache compression methods hit roughly 2-4× compression before quality degrades meaningfully. VFold's contribution is demonstrating that value cache merging adds another ~2× on top of existing methods, pushing toward 4-8× total compression. The next concrete target is enabling 128K+ context windows on consumer GPUs (24GB VRAM) for 70B-parameter models — currently infeasible without aggressive compression. If VFold's composability holds at that scale, it unlocks long-context inference without A100-class hardware. The obvious experiment not run: scaling to truly long contexts (128K-1M tokens) where the KV cache dominance is most extreme and where the inter-layer similarity assumption might break down as the cache grows. My read: probably (a) — they ran out of compute. Testing at 128K+ contexts with 70B models requires substantial resources, and the paper's contribution is the insight and method, not the scaling proof. The next paper from this group will likely be the scaling study.