MIT and Tel Aviv's LSL Stops AI from Forgetting Skills After Fine-Tuning

Researchers from Tel Aviv University and MIT propose Local Support Learning, a gating-based method that lets LLMs learn new skills without erasing old ones.

·
·
·
MIT and Tel Aviv's LSL Stops AI from Forgetting Skills After Fine-TuningPRO
  • Tel Aviv University and MIT propose Local Support Learning to stop catastrophic forgetting in LLMs.
  • LSL pairs a LoRA-style adapter with a GMM gate that fires only on in-distribution activations.
  • Retains 99% of pretrained capability on 7B models while learning new tasks at full strength.
  • Baseline LoRA drops 44% after one finetuning phase; LSL stays flat across three phases.
  • Decouples learning rate, batch size, and rank from forgetting, removing a key tuning headache.
  • Code and website at assafbk.github.io/lsl; overhead comparable to standard LoRA.

Fine-tuning a pretrained model on chemistry can improve domain performance while eroding coding, mathematics, or instruction-following skills. This failure mode, called catastrophic forgetting, complicates sequential adaptation because teams may lack the original pretraining data needed for replay. A Tel Aviv University and MIT preprint, Local Support Learning (LSL), proposes gated adapters that confine each update to activation patterns resembling the current fine-tuning set.

For a weight matrix W and input activation x, an update ΔW changes the layer output by ΔWx. The output remains unchanged only when x falls in the update’s null space, the set of vectors that ΔW maps to zero. Ordinary gradient updates therefore alter responses for many activation patterns, including those associated with prior capabilities. LSL learns where each update should apply and suppresses it elsewhere.

A gate for every update

Each fine-tuning phase pairs a standard adapter, such as LoRA, with a gate learned from current-phase activations. The adapter optimizes the task loss. The gate estimates whether a new activation lies within the local support of its training distribution. Conceptually, a gated layer computes Wx + g(x)ΔWx, where g(x) reduces the adapter’s contribution outside that support.

Component Role Practical effect
LoRA-style adapter Learns the task-specific weight update Preserves the parameter efficiency of adapter tuning
Density gate Models current-task activations Rejects inputs far from the fitted distribution
Per-matrix placement Routes updates using each layer’s activations

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads