Researchers from Meta, MIT and the University of Washington have tested language models that manage their own working context as an editable file. The model can preserve key facts, remove irrelevant steps, record failed experiments and maintain progress notes instead of relying entirely on fixed summarization, compression or retrieval rules.

The team evaluated zero-shot instructions, reusable context-management skills and reinforcement learning. On BrowseComp-Plus, a zero-shot Context Language Model raised accuracy by 11.4% while using 21.5% fewer floating-point operations. A reinforcement-learned Qwen3.5-9B improved from 28.8% to 42.5% on the same benchmark while using 12% fewer operations.

Results varied by task. The researchers also reported lower compute and higher scores on EdgeBench and larger gains in selected ContextBench tests. These are benchmark outcomes under specified conditions, not evidence that every long-running agent will become cheaper or more accurate.

Editable context creates its own failure modes. A model can delete information it later needs, preserve a prompt injection or turn a self-generated instruction into a persistent rule. Rewriting also complicates caching. The approach makes context management learnable, but production systems still need safeguards and external records for facts that an agent must not silently alter or forget.