New report: The state of agentic finance 2026
Valyu
Paper passages arranged and revised on a drafting desk, illustrating an agent editing its working context.
Research

Context Language Models: How Agents Edit Their Own Context

>_ Harvey Yorke

A research agent reads a filing, follows a reference and looks up a competitor. A few dozen steps later, useful evidence sits alongside abandoned searches, repeated passages and old plans.

What should it keep for the next step?

Context Language Models (CLMs) let the model make that decision by editing its own live context as a file. It can delete, compress and reorganise earlier turns. The edited file becomes the context for the next model call.

The Context Language Models paper, released on 29 September 2026 by Rulin Shao and colleagues, tests this approach on research, coding and extended optimisation tasks. For anyone building research agents, it opens up a useful question: how much control should an agent have over its own working memory?

Based on version 1 of the paper. Checked 2 October 2026.

What is a Context Language Model?

A Context Language Model is a language model that can edit the information supplied to its next call. In the paper's implementation, the current conversation is mirrored into a file. The model can rewrite that file using code or shell commands, and the agent runtime loads the changes for the next turn.

That synchronisation is the defining detail. A notes file only changes the model's next input if the runtime reads it back into the request. With a CLM, editing the designated context file directly changes that input.

The approach works with existing models without initial retraining. The researchers also explore improving it through reusable instructions and reinforcement learning. The model's context window stays the same size; the agent chooses what occupies it. Paper, section 4.1.

The official implementation provides a starting point for experimenting with the approach.

Title and abstract from Context Language Models, arXiv:2609.37725v1.

How do Context Language Models work?

The loop has four steps:

  1. Mirror the live context. The runtime writes the current conversation to a file and tells the agent where it is.
  2. Let the agent edit. It can remove failed searches, shorten tool outputs, update findings or retain exact passages it expects to need.
  3. Load the revised context. The runtime uses the edited representation for the next model call.
  4. Continue the task. New turns append normally until the agent chooses to edit again.

The edit affects the next model request. Illustration of the mechanism described in section 4.1 of the paper.

An agent can replace ten unsuccessful searches with a short record of what it already tried. It can maintain a table of unresolved questions, update a finding when new evidence arrives or keep an exact passage it expects to cite.

The paper observes agents writing loops to shorten search results, preserving document references, creating internal notes and reusing their own compaction helpers. In one multi-agent example, an agent makes 163 in-place updates to a tracker while keeping context around 6,000 to 8,000 tokens. Paper, figure 3.

What the paper shows

The most relevant result for research builders is BrowseComp-Plus: a benchmark of difficult questions answered using search and document-reading tools.

EvaluationAnswer or task qualityCompute
BrowseComp-Plus59.4% accuracy; 11.4% relative improvement over the strongest baseline21.5% fewer prefix-reuse FLOPs
EdgeBench-10About 5% higher scores than the summary baseline59% fewer prefix-reuse FLOPs

These are the authors' results with Qwen3.6-27B and a 32K context limit. BrowseComp-Plus uses a fixed research corpus; EdgeBench-10 covers ten repository tasks with 12-hour runs. Each row compares against its own baseline. The 11.4% improvement is relative, rather than 11.4 percentage points. Paper, section 5.1 and appendix E.

FLOPs measure computation. They do not directly measure your API bill or latency. The authors' prefix-reuse metric includes processing new input, generating output and reprocessing input after an edit breaks a cached prefix.

The model still needs to make good editing decisions. In the reinforcement-learning experiment, the smaller Qwen3.5-9B CLM initially trails the summary baseline. Training improves it from 28.8% to 42.5% accuracy on BrowseComp-Plus. Edit access alone is no guarantee of better answers. Paper, section 5.2.

What happens to prompt caching?

Prompt caching saves computation by reusing text the model has already processed. It usually depends on the beginning of the input staying unchanged. Edit that prefix, and the serving system may need to process the remaining text again.

That concern appears in the reactions to the paper. Alex Zhang's response on X welcomes the direction while questioning the cache implications. Nathan Lambert's announcement emphasises the freedom to move beyond rigid context-management policies. The practical tension is flexibility versus the cost of making edits.

The paper accounts for reprocessing in its standard-serving measurements. It also introduces Suffix Cache Reuse, which preserves cached states for surviving text beyond the unchanged prefix. Those states retain information from the earlier prefix, so this is a serving technique with an accuracy tradeoff to test. In the reported experiment, it reduces server computation by 35% while matching measured task performance. Paper, sections 4.3 and 5.3.

Changing a prompt or writing a context file does not enable Suffix Cache Reuse on a hosted model API. That requires support in the serving system. Measure input tokens, cached tokens, output tokens, latency and task quality on the stack you actually use.

CLMs, RAG, compaction and recursive language models

These approaches operate on different parts of an agent's information flow:

ApproachMain jobWhat the application must still decide
Retrieval-augmented generation (RAG)Retrieve relevant material from an external collectionHow retrieved evidence enters and leaves working context
CompactionReplace accumulated history with a shorter representationWhen to compact and which details must survive
Recursive Language Models (RLMs)Inspect long inputs programmatically and make recursive model calls over selected partsHow the live conversation is managed as work proceeds
Context Language Models (CLMs)Let the model rewrite its live contextHow to preserve source evidence and evaluate the edits

The RLM paper treats long input as an external environment the model can inspect. The CLM paper focuses on editing the live context itself. Retrieval, recursive inspection and editable context can therefore be combined in one system.

For a research agent, search determines which evidence it can access. Context management determines which of that evidence it carries into its next decision.

What should a research agent remember?

Consider an agent answering: "Has pricing power weakened across these portfolio companies?"

It reads management commentary, compares competitors and checks financial disclosures. Its working context might hold the current hypothesis, supporting findings, contradictions and questions still to investigate.

We would pair editable context with a separate evidence store. Save the retrieved passages before compressing them. Give each stored item an identifier, and retain its source URL, retrieval time and available document metadata.

A proposed architecture for research applications: context edits update working notes; the evidence store preserves the material behind them. This is our application design, not a benchmark result from the paper.

An agent can then shorten a long tool result to a compact record such as:

JSON
{
"question": "Has pricing power weakened?",
"status": "unresolved",
"findings": [
{
"observation": "Management describes promotional pressure.",
"evidence_ids": ["E017"]
}
],
"open_questions": [
"Does the margin change reflect pricing, costs or product mix?"
]
}

This is an illustrative ledger. E017 stands for a stored passage; the observation is a fictional example.

Before writing the final claim, the agent loads E017 and checks its original wording, date and scope. A summary may be enough to choose the next search. The supporting passage is what lets someone review the conclusion.

The same pattern applies to scientific research. A pipeline review needs to retain trial identifiers, phases, endpoints and conflicting findings even as it compresses the reading history. Preserve facts that would change the answer, including evidence against the current hypothesis.

Where Valyu fits

Valyu provides the search layer for this architecture. An agent can ask a natural-language question across web, academic and specialist sources, then retain the returned content and source URLs in its own evidence store. See the Search API documentation.

For example, the retrieval step can ask for a company's relevant filing passages directly:

Python
from valyu import Valyu
 
client = Valyu()
response = client.search(
"Apple 10-K disclosures about pricing pressure and gross margins",
included_sources=["valyu/valyu-sec-filings"],
)
 
if not response.success:
raise RuntimeError(response.error or "Search failed")
 
for result in response.results:
print(result.title, result.url)
print(result.content)

Install the Python SDK and set VALYU_API_KEY in your environment. This snippet demonstrates retrieval. Your runtime implements evidence storage and context editing; inspect the returned passages before using them to support a claim.

Our SEC filing indexes let agents request relevant content by company, section and subject, without supplying filing identifiers first. The same interface reaches sources such as academic papers, clinical trials and patents. The catalog documents availability by plan and access mode.

If you want a completed research task, Valyu DeepResearch handles planning, search and synthesis and returns a cited report. You can use that report and its sources as input to a larger agent workflow.

If you're an agent: Set up Valyu and make your first search. The instructions cover authentication and getting started.

If you're a human: Create an account and get an API key. Connect through a Python SDK or MCP.

What to test before adopting editable context

Start with a research task that already has a measurable outcome. Run your current context policy and a version with model-directed edits against the same questions and sources.

Track four things:

  • Answer quality. Did the agent resolve the question, including material contradictions?
  • Evidence retention. Can it still recover the passage behind every consequential claim after editing?
  • Efficiency. What happened to total tokens, cache use, latency and actual cost?
  • Failure recovery. Can you restore a previous context version after an edit removes something essential?

Keep versioned context snapshots. In an application, protect system instructions and permissions outside the editable working state, and treat retrieved content as evidence rather than instructions. The paper itself identifies editable context as a potential way for prompt injections to persist. Paper, section 6.

Common questions

Can an existing language model become a CLM?

Yes. The paper evaluates existing models with a runtime that exposes editable live context. Further training can improve the behaviour, but initial retraining is not required. Writing an ordinary notes file is only part of the mechanism: the runtime must use the edits in subsequent model requests.

Do CLMs give agents unlimited memory?

The active context window remains finite. Editing gives the agent control over its contents. Recovering removed information requires an external copy or another retrieval step.

Do Context Language Models replace RAG or search?

No. CLMs manage the active context. RAG retrieves information from an external collection, and search discovers evidence the agent has not yet read. These capabilities can work together. For auditable research, retain the source material behind compact working notes.

Give the agent room to think

CLMs let an agent choose what it carries forward as a task develops. For research, a useful working context holds the question, current findings and unresolved contradictions. The evidence behind those findings should stay recoverable through every edit.

If you're an agent: Follow the Valyu onboarding instructions.

If you're a human: Create an account and get an API key, or run your first DeepResearch task.

Valyu Add

Join 12,000+ professionals and knowledge workers.

Valyu Add is a free weekly research briefing for builders, investors and operators. Every issue is sourced, cited and verified with Valyu DeepResearch.