MIT and Tel Aviv's LSL Stops AI from Forgetting Skills After Fine-Tuning
Researchers from Tel Aviv University and MIT propose Local Support Learning, a gating-based method that lets LLMs learn new skills without erasing old ones.
- Tel Aviv University and MIT propose Local Support Learning to stop catastrophic forgetting in LLMs.
- LSL pairs a LoRA-style adapter with a GMM gate that fires only on in-distribution activations.
- Retains 99% of pretrained capability on 7B models while learning new tasks at full strength.
- Baseline LoRA drops 44% after one finetuning phase; LSL stays flat across three phases.
- Decouples learning rate, batch size, and rank from forgetting, removing a key tuning headache.
- Code and website at assafbk.github.io/lsl; overhead comparable to standard LoRA.
Fine-tuning a pretrained model on chemistry can improve domain performance while eroding coding, mathematics, or instruction-following skills. This failure mode, called catastrophic forgetting, complicates sequential adaptation because teams may lack the original pretraining data needed for replay. A Tel Aviv University and MIT preprint, Local Support Learning (LSL), proposes gated adapters that confine each update to activation patterns resembling the current fine-tuning set.
For a weight matrix W and input activation x, an update ΔW changes the layer output by ΔWx. The output remains unchanged only when x falls in the update’s null space, the set of vectors that ΔW maps to zero. Ordinary gradient updates therefore alter responses for many activation patterns, including those associated with prior capabilities. LSL learns where each update should apply and suppresses it elsewhere.
A gate for every update
Each fine-tuning phase pairs a standard adapter, such as LoRA, with a gate learned from current-phase activations. The adapter optimizes the task loss. The gate estimates whether a new activation lies within the local support of its training distribution. Conceptually, a gated layer computes Wx + g(x)ΔWx, where g(x) reduces the adapter’s contribution outside that support.
| Component | Role | Practical effect |
|---|---|---|
| LoRA-style adapter | Learns the task-specific weight update | Preserves the parameter efficiency of adapter tuning |
| Density gate | Models current-task activations | Rejects inputs far from the fitted distribution |
| Per-matrix placement | Routes updates using each layer’s activations |
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.