We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file.…

45 points•emersonmacro•about 4 hours ago•10 comments•

10 comments

bob102931 minutes ago
I would be concerned with context management consuming limited attention resources.

Do you want your agent solving its own memory crisis, or do you want it solving the actual task? It can probably do both at the same time, but I suspect there is a non trivial cost associated with this.

A separate hypervisor agent that manages the main agent's context would be much better in my experience. You can run it on a different schedule and the main agent has to spend zero tokens thinking about it. This also makes it a lot easier to control when caches will be missed.

alightsoul21 minutes ago
Maybe have A second model do the management?
theroadnotbacon15 minutes ago
Looks like they tested that in the paper, and the code allows for it as well. https://github.com/facebookresearch/context-language-models/... Of course, this is for suggesting context management strategies rather than the actual management afaict.
svachalekabout 1 hour ago
Wow. Context management is one of the big remaining hassles with modern LLMs so this could be big. The obvious complication is cache busting so it's also exciting they investigated solutions for that.
Bolwin35 minutes ago
The biggest discovery might actually be that they ignored regular caching rules and kept invalid cache suffixes and it didn't hurt performance
visarga35 minutes ago
Can't we do this trick today with any model? Just send the file as next context. Of course you pay the price for cache misses, depending how deep you make changes, while CLM just ignores the recomputation.
nsingh232 minutes ago
One approximation of this is the experimental context management Codex has been moving towards (not released yet). Rather than relying on summary compaction, the model maintains notes as it works and as it approaches the context limit. A new session is just a fresh context with those notes attached, and a pointer back to the previous session.

Not exactly like what this paper is suggesting, but similar in the sense it lets the model decide what and how to persist across turns.

I recreated this in Pi, with a max token limit on how long the note can be, to pressure the model to be concise. Ends up being cheaper than summary compaction too.

ijidak2 minutes ago
Memento
visarga28 minutes ago
That is similar to what I am thinking... not just edit the context as a file or string, but have a way to evict blocks and replace them with summary notes and also be able to retrieve them on demand.

Do you have a public repo for your approach?

Read the full thread on Hacker News →

Related stories