The first frontier LLM with ten million tokens of context.
1 comment
This family of LLMs has ten million tokens of native context, but it's really a proof of concept for a new kind of attention, called VSA (Voltropy Scalable Attention.)
There were previously two approaches to getting context to work at this scale. One was to approximate attention in ways that degraded reasoning-quality at lower context lengths. The other was to abandon the transformer architecture in favor of cheaper alternatives that only did e.g. recurrent computation, like state space models.
We designed VSA to be a third path: One that would extend context by an order of magnitude while also increasing reasoning quality at shorter contexts. And what is striking about the results we share today is that this worked, even when we retrofitted VSA to transformers that were training with conventional attention.
When we augmented DeepSeek v4.0 Flash with VSA, the resulting model scored better at ten million tokens than the original did at one million tokens. To our surprise, it did even better than DeepSeek's next-generation model, v4.1 Flash. In fact, the 1M-token performance was so strong that it beat Anthropic's Fable 5.1 by around 10%.
All this is to say that we believe we've found an important new primitive for building LLMs. As we discuss in the blog post, we think that this could scale way beyond 10M tokens, and that stretching it to billions or trillions of tokens could fundamentally reshape what LLMs are.
First, it could solve continual learning, by making context windows so large that agents can learn continuously within ever leaving a context window. Basically, a loophole to solve continual learning without post-deployment weight updated.
Second, it could allow for much of the data that is currently stored in weights through lossy memorization to be moved into a structured, context-like format. Basically, the goal would be to make LLMs less like human brains and more like classical Von Neumann machines.
Obviously, that vision requires scaling VSA far beyond where it is today. But we're excited to share this first step on the journey with you all.
Read the full thread on Hacker News →
Related stories
- You Don't Need a %Frontier LLM%rakshazi.meHacker News · 1 points · 8 days ago
- Hacker News · 1 points · 3 days ago
- Hacker News · 1 points · 2 days ago
- Hacker News · 1 points · 5 days ago
- Hacker News · 1 points · 10 days ago
- DEV Community · 8 points · 20 days ago