The standard contract between a language model and its context is append-only. Every message, tool response, and reasoning step accumulates. Nothing leaves. A model can summarize things into a smaller string, but the summary replaces the original only if an external harness decides to do that swap — the model itself has no lever.

A new paper from Meta AI and Allen AI breaks that contract. Context Language Models (CLMs) treat the working context as an editable file. The model gets a path to it in the system prompt, and from there it can use ordinary Bash utilities — sed, Python, file operations — to read, edit, reorder, or discard whatever it’s carrying. When no edit happens, tokens append as normal. When an edit does happen, the changes take effect on the next inference call.

The practical consequence is that models can become deliberate about what they carry forward. On a long research task, instead of accumulating every discarded search thread, a CLM can compress completed branches and delete working notes it no longer needs. This is not just tidier — it’s cheaper. Every subsequent inference call processes a shorter context, paying fewer FLOPs for the same effective information.

The numbers are striking. On EdgeBench’s 12-hour tasks, CLMs scored 5% higher than summary-based approaches while consuming 59% fewer FLOPs. On BrowseComp-Plus (a deep research benchmark), the improvement was 11.4% accuracy with 21.5% fewer compute operations. On 24-hour multi-agent software tasks, CLMs delivered 65% greater downstream speedup versus baseline agents operating identically.

What the paper emphasizes is that the models aren’t using specialized context management primitives — they’re using the same tools a programmer would reach for. The paper reports models spontaneously developing habits like using Python regex for surgical edits, offloading reference material to disk during a task, and defining helper functions in the context file to avoid restating them each turn. None of this was explicitly trained; the models figured it out.

There’s also an RL result worth noting: online reinforcement learning on top of the CLM formulation improved Qwen3.5-9B on BrowseComp-Plus by 47.6% while using 12% fewer FLOPs compared to RL on a standard LM. The CLM setup gives RL a better credit-assignment surface — the model’s context edits are explicit signals about what it judged worth keeping.

The serving side gets a corresponding optimization called Suffix Cache Reuse (SCR). Normally, when a model edits something mid-sequence, the KV cache states for every token after the edit point become stale and must be recomputed. SCR instead reuses the cached states for tokens that survived the edit, adjusting rotary position encodings to reflect new positions, and only re-computes what actually changed. On BrowseComp-Plus, SCR matches standard SGLang performance at 65% of the compute cost.

The multi-agent extension is straightforward in this framing: since the context is a file, multiple agents can share synchronized context files naturally. The paper demonstrates this on the 24-hour software benchmark.

The broader shift here is architectural. Context management has mostly been an external concern — the harness decides what goes in the prompt, tools summarize completed work, retrieval layers inject relevant chunks. CLMs absorb that layer into the model itself. Whether that’s better depends on how much you trust the model’s editorial judgment. The evidence here suggests the trust is justified, at least at the scales tested: the model editing its own notes turns out to produce better outcomes at lower cost than handing a fixed summary to the next call.

The code is at github.com/facebookresearch/context-language-models.