931 comments
Data at https://gertlabs.com/rankings
This is a plus in my opinion.
This is the actual big announcement. 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.
Cache doesn't help you much when you are compacting every 5 minutes...
I was shocked at how quickly I ran out my $100/mo subscription with a single agent (sol medium).
No LLM will be cost effective if it's compacting this often. You have to find a way around it.
The longer your chat gets, the slower and more expensive it gets.
Subagents are expensive but they scale way closer to O(n) than O(n^2).
Have some agents make bug reports/feature requests/roadmaps (linear is very AI friendly), others coordinate, others work on grinding out an individual ticket.
If there is a good ticket-level description, it's a waste of time IMO to have a main agent do it, that should be an agent with fresh context that will do it better faster (the shorter the context, the better models are at using the context they're given).
It's crazy on Codex. I sometimes get just 2-3 turns before it compacts. It has forced me to use persistent project documentation for everything. Maybe that's not a bad thing but unless it reads all the documentation after every compaction (and uses half its cache), it goes off the rails. By comparison, Opus 5.5 is a breath of fresh air. It takes FAR longer to hit the cache limit and that means it keeps useful information in working memory far longer. I think this alone has resulted in a massive productivity and efficiency increase for me.
Suggest trying it out yourself: Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does. The difference is stark.
Edit: Defining "difficult" as a complex coding or systems task (or even series of them in a single prompt).
Also, a lot of this work is verification to ensure that AI generated code does what is intended and is safe to merge and deploy. That verification work is critical and uses a lot of tokens.
My main worry is that we get to a point where they have something much, much smarter than anything public and access is gated by extraordinarily high costs.
That probably can't happen though right? Inference is surprisingly cheap compared to the training.
This model might be the first step in that direction, as competition heats up between OpenAI and Anthropic.
It's also used to describe the SFT bootstrapping for posttraining, which is what people generally refer to as Chinese labs "distilling".
I would almost guarantee that smaller US frontier models (ex Luna/Sonnet) are distilled from their respective large models.
Have they, actually? A lots of speculation and claims but what is the level of admittance?
- Training data, harnesses, and processes, that drives the quality of the models. Several companies have those. Some of those companies are in China.
- Infrastructure and funding for running those. That's a scarce commodity currently, training the latest frontier models cost billions of dollars apparently. And if you don't have the infrastructure already and don't have the suppliers on speed dial, good luck getting anything.
- Infrastructure for running inference for running what comes out of those. Several of the key providers of this infrastructure are using their own in house chips for this now. At scale this means huge cost savings.
If you start from scratch without infrastructure, there are a bunch of open source things you can find. But beyond that, you'll have a lot of catching up to do. That's the very definition of a moat.
Infrastructure is readily available. Anthropic and OpenAI don’t own anything, they’re just paying for compute. You might not be able to buy 100k GPUs right now but you can rent it.
You only need to look at how Jev immediately became one of the highest volume models after launch on OpenRouter to see how this industry is moatless.
OpenAI and Anthropic employees frequently leave to start their own labs and raise hundreds of millions for it which they can use to immediately pay for infrastructure and training data. At most OpenAI and Anthropic have… brand and talent.
If anything, OpenAI and Anthropic are heavily disadvantaged because they have huge long term financial commitments that have backed them into a corner, likewise the regulatory pressure… startups have none of that.
Which, honestly, is fine. A lot of juice to squeeze in efficiency and even if models got zero more capable, making the capability that is already here cheaper is a huge win for everyone (except Nvidia)
Not once has any of these predictions come true, the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention.
So why now? What is special about today that suggests all of this is coming to a screeching halt despite all evidence to the contrary?
It just seems like these claims are constant and looking back the calls of 'plateau' between 2023 and 2025 were clearly false, why should we think it's different now?
It's not even anything controversial..
it's also why there have been so many calls for regulation and slowdowns.
I see posts about OpenAI and Anthropic latest and don’t even care looking at what they do better. I just read the comments here.
I use DS4.1 Flash and GLM 5.3 Flash, pay peanuts per day and get more than acceptable results.
Insane pricing pressure on the horizon. Even if big companies will not go with open weight models, the threat will be ever present that they can instantly flip flop on providers.
I remember when bandwidth was super expensive and now it’s dirt cheap.
Consumers are now saying the new pricing with lower usage caps is not so great. https://news.ycombinator.com/item?id=49896975
Sol 6 was so bad that I switched over to Opus 5.5 exclusively.
Huge regression compared to Sol 5.6, often doing really dumb things. Same for Luna.
Even Astra is very unreliable for coding. Brilliant for vision, sometimes just great, but it also often does very stupid things.
I'm a bit sour on OpenAI right now and skeptical that 6.1 will be much different.
(Note: this is after preferring and shilling Codex/OpenAI models for the last half year)
I've implemented multiple features side by side with Opus 5.5 and 6 Sol, and the Opus 5.5 results always have fewer high severity bugs and require fewer rounds of fixes to get it over the finish line.
If 6.1 Sol has actually matched Opus 5.5, I'd be very happy. However, benchmarks and real usage don't seem to agree in my own tests. So we'll have to see.
For coding specifically, I've found 5.6-Sol > 6.0 Sol > Astra.
For modeling and artwork, Astra has been great routinely outperforming Kimi.
This is reminiscent to me of what Anthropic pulled back in February with their adaptive thinking rollout.
I can't wait for technology to catch up to a point where we can rid ourselves of this oligopoly.
I had to switch back to 5.6 Sol after trialling 6 Sol for like 3 days - I was getting insanely annoyed at how misaligned it is. Will try 6.1 but not very high hopes
And, is it really even an oligopoly anymore? Open weight models are incredibly competitive in every way; whether you want to use US providers, Chinese official providers, self host, etc.
I'll just leave this here: https://marginlab.ai/trackers/codex/
I tried out fable 5.1 the day it was released and coming from gpt-5.6-sol I was truly mind blown (both in terms of code and prose it was generating - outputs I could finally enjoy reading and looking at).
Then when opus 5.5 came out, again same thing + far cheaper and faster.
I went from using OAI exclusively the entire year to a point now where i haven’t touched one of their models in at least a few weeks now.
I think OAI has lost the plot. OAI models simplify have no taste. And I don’t mean in front-end design way (although that too). They have no taste in how the model writes code, how it writes prose, how it writes in-line comments, how it writes documentation, or how it even picks variable names. There’s just no taste throughout.
Anthropic models are very thoughtful and have so much taste all around.
And, of course, GPT-6 came out as Anthropic fixed a bunch of stuff with their models -- faster (via fewer tokens, and TPS for Sonnet), easier to work with, better results, cheaper (via pricing and, again, fewer tokens). I don't know if the timing and the suddenness of the improvement on Anthropic's side sharpened the vibes comparison this round, but Internet opinion went pretty clearly to Anthropic.
FrontierCode's results make it look like Sol-6.1 may slot in well where you'd use Sonnet or Opus's low effort.
One thing I don't think any of this reflects is that many well-specified coding tasks, including the self-testing and doing research and tracing out dependencies and so on, aren't really bleeding-edge now: Luna-5.6 and small open models handle them fine. Stuff like "why is this box dropping connections?" or "here's a thing I want you to model/figure out" can benefit from bigger models. But far from everything does!
Read the full thread on Hacker News →
Related stories
- GPT-6.1 Sol replaces GPT-6 Sol after just 7 days, with near-Astra intelligenceartificialanalysis.aiHacker News · 47 points · about 16 hours ago
- GPT-6.1 Sol: Release Intelligence, Performance and Priceartificialanalysis.aiHacker News · 2 points · 1 day ago
- Addendum to GPT-6 Astra System Card: GPT-6.1 Soldeploymentsafety.openai.comHacker News · 3 points · 1 day ago
- Hacker News · 4 points · 6 days ago
- GPT-6 Sol (Max) Intelligence, Performance and Price Analysisartificialanalysis.aiHacker News · 2 points · 8 days ago
- Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price warsimonwillison.netHacker News · 2 points · 8 days ago