Imagine a hotel where every room has its own doorman. Each doorman only recognizes guests who checked in during their shift — and politely ignores everyone else. When a new shift starts, the new doorman learns only the new guests, but crucially, the old doorman is still standing there, still recognizing his own. Nobody gets locked out. That's Local Support Learning (LSL) in a sentence: per-phase gating that activates weight updates only for inputs that look like the data from that phase. The committed claim is straightforward but important: you can resolve catastrophic forgetting in LLMs up to 7B parameters by pairing a standard weight adapter (LoRA or similar) with a Gaussian Mixture Model gate that fires only when the input activation falls within the training distribution of that adapter. The gate's likelihood decays rapidly away from its training data, so it naturally stays closed for inputs from earlier or later learning phases. No prior data is needed. No replay buffer. No knowledge distillation from a frozen teacher. The update is geometrically local — it only touches the part of activation space where the new data lives. The architectural insight that drives LSL is geometric: the authors frame forgetting not as a loss-landscape problem but as an input-space problem at each weight matrix. Gradient-based optimizers minimize the training loss but are agnostic to what happens to activations from prior distributions — they'll cheerfully overwrite old features if that helps the new objective. LSL's fix is to treat each adapter as a conditional module: it fires only when the GMM gate says 'this looks like mine.' The GMM is a natural fit because its probability density drops off exponentially outside its fitted clusters, giving you a cheap, well-calibrated switch without learning a separate binary classifier. On the ladder, LSL is positioned against the continual learning literature — methods like EWC (Elastic Weight Consolidation), PackNet, replay-based approaches, and parameter-efficient finetuning baselines like vanilla LoRA. The key comparison is retention of pretrained knowledge (measured on general benchmarks) while maintaining finetuned performance across sequential task phases. The paper reports that LSL retains both pretrained and finetuned capabilities, while standard LoRA and even some continual learning baselines degrade substantially as phases accumulate. The approach scales to 7B parameters, which places it in the territory that matters for practical LLM deployment. However, the strongest continual-learning baselines (like O-LoRA and certain replay methods) are not always dramatically outperformed — the margin depends on the scenario. Integrity is reasonable but has the usual self-graded-homework caveats. The evaluation spans multiple training phases and uses standard LLM benchmarks, which is good. The authors report robustness to hyperparameters, which is a strong practical signal. But there's no independent replication, no pre-registration, and the benchmark suite — while reasonable — was presumably selected by the authors. The GMM gate adds minimal overhead (the paper emphasizes memory and compute efficiency), but the exact overhead numbers at scale deserve scrutiny. Code and a project website are provided, which lowers the replication barrier. The milestone to watch is whether this approach holds across 10+ sequential finetuning phases and across heterogeneous task types (not just instruction tuning variants). The current demonstration is compelling at a handful of phases; the real stress test is whether GMM gates remain discriminative as the number of phases grows and activation distributions start to overlap. If LSL can handle 20+ phases on 70B+ models without gate confusion, it becomes a serious infrastructure component for lifelong-learning LLM deployments. The obvious experiment the authors did not run: testing on 70B-scale models and on truly diverse sequential tasks (e.g., alternating between code, multilingual, and domain-specific finetuning). My read is (a) compute budget — running 70B experiments with multiple sequential phases is expensive and this is a four-author academic paper, not a lab with 10,000 GPUs. The framework is general enough that nothing should break at that scale, but 'should' and 'does' are different words.