Meta Lets CLMs Rewrite Their Own Memory, Cutting AI Compute by 35%

A new paradigm lets language models edit their own context file with arbitrary Bash commands, beating hand-designed summarization harnesses across long-horizon agent tasks.

·
·
Meta Lets CLMs Rewrite Their Own Memory, Cutting AI Compute by 35%PRO
  • CLMs treat context as an editable file the model rewrites via Bash, replacing harness-controlled compaction
  • Zero-shot beats Codex-style summary by 11.4% on BrowseComp-Plus using 21.5% fewer FLOPs
  • On 24-hour multi-repo agent swarm, CLM delivers 65% greater downstream speedup at matched compute
  • Success-gated GRPO lifts Qwen3.5-9B by 47.6% on BrowseComp-Plus with 12% fewer FLOPs
  • Suffix Cache Reuse patches SGLang to reuse cache after mid-context edits, cutting server compute 35%
  • Code and Pi-agent integration available at facebookresearch/context-language-models

Long-horizon agents still delegate context management to hand-built harnesses. When a conversation approaches its token limit, application code decides what to summarize, move elsewhere, or discard. Researchers at the University of Washington and Meta Superintelligence Labs propose giving those decisions to the model itself.

Their paper introduces Context Language Models, or CLMs. A CLM treats its live context as an editable file, modifies that file with Bash commands, and synchronizes the result with the inference server. In the authors’ tests, existing models using this design surpassed every evaluated context-management harness without additional training.

Fixed compaction drops state

Long-running agents commonly trigger compaction at a predefined context length and apply a fixed summarization procedure. Cursor, Codex, and Terminus2 use variants of this approach, while newer systems let models choose from constrained actions such as compact, offload, or retrieve. Each design limits how the model can reorganize its working state.

ContextBench separates context management from general knowledge and reasoning through four synthetic tasks:

  • Needle Retention: Preserve specified lines verbatim across chunks of filler text.
  • Sudoku Sketchpad: Maintain a 16 × 16 board as users submit one move at a time.
  • KV Store: Offload batches of key-value writes and retrieve exact values later.
  • Log Triage: Store log lines and answer lookup or count queries.
ContextBench tasks and context-management results
ContextBench tests exact retention, mutable state, retrieval, and offloaded log analysis.

These tasks expose different weaknesses in fixed harnesses. Summary-based compaction can omit or invent details during Needle Retention and Sudoku Sketchpad. Methods lacking in-place editing must regenerate the full Sudoku board after each move. Standard coding tools can offload information, yet often cannot remove selected material from the live context. Every evaluated baseline fails at least one of the four tasks.

A file becomes working memory

The CLM implementation mirrors the live context to a file and supplies its path in the system prompt. The model can inspect or edit that file with shell commands, after which the harness synchronizes the contents before the next inference request. Unedited contexts continue to append tokens normally.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads