1050 points•crorella•1 day ago•931 comments•

931 comments

gertlabsabout 10 hours ago
We tested it across 100 unsaturated coding and engineering environments. Both Astra and 6.1-Sol are pretty comfortably ahead of Opus 5.5 in these types of evaluations, and both end up being cheaper than Opus via API usage. 6.1-Sol is also cheaper than Sonnet 5.5 and much smarter. The only verifiable domain Anthropic seems to be clearly ahead is chemistry (and perhaps also some unverifiable domains like being pleasant to work with, since GPT-6 models have a tendency to be low initiative beyond what the prompt tells them). Highly recommend using the OpenAI Flex endpoint for any API work.

Data at https://gertlabs.com/rankings

wpmabout 10 hours ago
"GPT-6 models have a tendency to be low initiative beyond what the prompt tells them"

This is a plus in my opinion.

jonaustinabout 5 hours ago
Maybe...at least in the web ui chats I now do several rounds of "carefully review your conclusions" because gpt is incredibly lazy and misses ridiculously obvious things (obvious with a little bit of web research) pretty much every single time it's asked to research anything.
dpc_01234about 5 hours ago
Nothing better than staffing your prompts/skills with long lists of all the things that you don't want, but agent might think you might want even if you didn't ask.
minimaxir1 day ago
> Cached input costs just $0.10 per million tokens—95% less than standard input pricing and 50% less than GPT‑6 Sol’s cached input pricing

This is the actual big announcement. 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.

joshstrange1 day ago
> 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.

Cache doesn't help you much when you are compacting every 5 minutes...

I was shocked at how quickly I ran out my $100/mo subscription with a single agent (sol medium).

redox991 day ago
If you run out of sol medium with $100 you're doing something wrong. Astra destroys your usage, I get 1 day of usage with Astra, but 6 sol is almost unlimited and I only use xhigh.
onlyrealcuzzo1 day ago
If you're compacting every 5 minutes, you have a workflow problem - period.

No LLM will be cost effective if it's compacting this often. You have to find a way around it.

jmalickiabout 12 hours ago
Use more subagents.

The longer your chat gets, the slower and more expensive it gets.

Subagents are expensive but they scale way closer to O(n) than O(n^2).

Have some agents make bug reports/feature requests/roadmaps (linear is very AI friendly), others coordinate, others work on grinding out an individual ticket.

If there is a good ticket-level description, it's a waste of time IMO to have a main agent do it, that should be an agent with fresh context that will do it better faster (the shorter the context, the better models are at using the context they're given).

Gareth321about 19 hours ago
> Cache doesn't help you much when you are compacting every 5 minutes...

It's crazy on Codex. I sometimes get just 2-3 turns before it compacts. It has forced me to use persistent project documentation for everything. Maybe that's not a bad thing but unless it reads all the documentation after every compaction (and uses half its cache), it goes off the rails. By comparison, Opus 5.5 is a breath of fresh air. It takes FAR longer to hit the cache limit and that means it keeps useful information in working memory far longer. I think this alone has resulted in a massive productivity and efficiency increase for me.

AmazingTurtle1 day ago
you can actually leverage 400k and 1M contexts in codex with very little code changes to the harness. note that excess context past the.. 250k or 400k mark (i don't remember) is charged at 2x the price.
TuxSH1 day ago
Exactly half as expensive as Opus 5.5 in every API pricing metric
bigwheels1 day ago
And half as good. I didn't have great experiences with Anthropic models in the past, but Opus 5.5 seems to have turned a major corner. It is churning through tasks significantly more quickly and efficiently.

Suggest trying it out yourself: Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does. The difference is stark.

Edit: Defining "difficult" as a complex coding or systems task (or even series of them in a single prompt).

dom961 day ago
Based on my benchmark[1] it is the same price as Opus 5.5 and just as capable.

1 - https://bench.killswitch-lang.org

pvab3about 12 hours ago
gpt 6 Sol was already supposedly better and 50% cheaper than 5.6 Sol right? I didn't understand why they were keeping 5.6 Sol
vcryanabout 13 hours ago
People's volume and approach varies. I'm a happy customer and I use my entire double max subscription on planning and analysis and have other models doing all my implementation work because I would burn through my subscription in a day or less. It's difficult to calculate, but I'm something like 10-20 billion token per week consumer and I can't use a US-based model to do this volume of implementation work.

Also, a lot of this work is verification to ensure that AI generated code does what is intended and is safe to merge and deploy. That verification work is critical and uses a lot of tokens.

verdverm1 day ago
cache is typically 10%, is this OAI setting a new level at half, 5%?
crazylogger1 day ago
The backdrop being deepseek offering 1% (I remember it was ~1% when 4-pro first came out early this year - 4-pro is now removed) / 2% (current for 4.1-flash).
I must say that this AI thing is going more or less as I felt it would back about a year ago. I think there is no real moat in AI models. It's a commodity and the big labs have predictably been caught in a race to the bottom. Not sure if this is going to turn better or worse for all of us common folks. I must say I'm a bit happy though in the sense that "intelligence" is not going to be controlled and be rented out by a small minority.
munksbeerabout 14 hours ago
Sort of. Hopefully the trend continues and we continue to get advances in the cheaper and open source models, but by all accounts, inside the frontier labs, they get to use much better models that they haven't released.

My main worry is that we get to a point where they have something much, much smarter than anything public and access is gated by extraordinarily high costs.

That probably can't happen though right? Inference is surprisingly cheap compared to the training.

wyreabout 10 hours ago
I don't think the limitation is the price of inference of extraordinary private models, but rather scaling to the demand of the model and if that's the case the labs would keep the models a secret, unless they need it for marketing purposes like we saw Anthropic marking Mythos.
jumploops1 day ago
The Chinese labs have shown that distillation is incredibly effective, but the major US frontier labs haven’t (yet) been incentivized to shrink their models in the same way.

This model might be the first step in that direction, as competition heats up between OpenAI and Anthropic.

ismael_rr1 day ago
Distillation is also a broad term - I think most specifically, it refers to training a smaller model on a larger/better model's full output token distribution rather than normal pretraining, which only can access the next token in the data that was actually used.

It's also used to describe the SFT bootstrapping for posttraining, which is what people generally refer to as Chinese labs "distilling".

I would almost guarantee that smaller US frontier models (ex Luna/Sonnet) are distilled from their respective large models.

nevir1 day ago
Isn't that roughly what Haiku/Sonnet/Luna/Terra are?
nicce1 day ago
> The Chinese labs have shown that distillation is incredibly effective, but the major US frontier labs haven’t (yet) been incentivized to shrink their models in the same way

Have they, actually? A lots of speculation and claims but what is the level of admittance?

ashleynabout 11 hours ago
If there is high upfront training cost (filtering, classifying data, etc) then what we may find is a situation similar to drug pricing where the cost of the end product remains high and the motivation to keep it closed is there to justify the research input. Any similarity to the pharma industry doesn't inspire confidence in cheapness and openness.
Asookaabout 9 hours ago
I don't think it will get quite that bad. Making your own pharmaceuticals is practically impossible, because you cannot buy the machines or chemicals needed without jumping over lots of regulatory hurdles. Training your own AI model is just a question of money and time. The hurdles there are mostly technical - how do you read the entire Internet without getting banned. Unless you envision a future where AI training itself is regulated.
jillesvangurpabout 12 hours ago
There are several moats:

- Training data, harnesses, and processes, that drives the quality of the models. Several companies have those. Some of those companies are in China.

- Infrastructure and funding for running those. That's a scarce commodity currently, training the latest frontier models cost billions of dollars apparently. And if you don't have the infrastructure already and don't have the suppliers on speed dial, good luck getting anything.

- Infrastructure for running inference for running what comes out of those. Several of the key providers of this infrastructure are using their own in house chips for this now. At scale this means huge cost savings.

If you start from scratch without infrastructure, there are a bunch of open source things you can find. But beyond that, you'll have a lot of catching up to do. That's the very definition of a moat.

reticulatesabout 10 hours ago
YC is pumping out training data startups, some of the fastest growing startups are training data startups because OpenAI and Anthropic and Meta and Google are paying them billions.

Infrastructure is readily available. Anthropic and OpenAI don’t own anything, they’re just paying for compute. You might not be able to buy 100k GPUs right now but you can rent it.

You only need to look at how Jev immediately became one of the highest volume models after launch on OpenRouter to see how this industry is moatless.

OpenAI and Anthropic employees frequently leave to start their own labs and raise hundreds of millions for it which they can use to immediately pay for infrastructure and training data. At most OpenAI and Anthropic have… brand and talent.

If anything, OpenAI and Anthropic are heavily disadvantaged because they have huge long term financial commitments that have backed them into a corner, likewise the regulatory pressure… startups have none of that.

rubslopesabout 23 hours ago
Yes. I have several subscriptions for cost saving and I'm all the time swapping models according to the cost/benefit necessary for the task at hand.
gradus_ad1 day ago
Ominous for the industry and investors that token price is becoming the main battleground. Could be Anthropic's rationale for IPOing this year.
mixdup1 day ago
Another piece of evidence on the pile that the sudden panic and desire to "slow down" is because they're hitting the plateau on capability

Which, honestly, is fine. A lot of juice to squeeze in efficiency and even if models got zero more capable, making the capability that is already here cheaper is a huge win for everyone (except Nvidia)

luma1 day ago
Some version of this claim has been made for the past 4 years. There's a data cliff, there's no more compute to buy, the financials don't make sense and all of these orgs will be out of business by end of quarter.

Not once has any of these predictions come true, the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention.

So why now? What is special about today that suggests all of this is coming to a screeching halt despite all evidence to the contrary?

CuriouslyC1 day ago
It's not so much that they're hitting a plateau in capability, as we're saturating long horizon benchmarks and it's not greatly improving general usability. On the other hand, newer models have been amazing for people interested in 3d, graphics, video editing, etc. The difference between Opus 5.5/Astra and earlier models is night and day even if for many coding tasks they're not a revolution.
sebzim45001 day ago
Is there anything that could happen that you wouldn't use as evidence that they are hitting a plateau?

It just seems like these claims are constant and looking back the calls of 'plateau' between 2023 and 2025 were clearly false, why should we think it's different now?

LPisGood1 day ago
Nvidia can start putting weights in silicon if model development slows down.
I think they are hitting compute restrictions. And buying compute right now can be 3-4X. And the costs are increasing. If they train a larger model and demand is high, that’s a lot of compute for Codex subscriptions, which is a loss leader for them. Especially Pro 20X which they just nerfed to 10X.
djfjkfkffkkf1 day ago
China will do to llms what they did to german cars
bogrollben1 day ago
I guess I'm out of touch. What did china do to german cars?
ok1234561 day ago
We can only hope.
Razengan1 day ago
Why make a new account just to post this comment?

It's not even anything controversial..

jorblumesea1 day ago
This is literally the plan, open weight models are something like 60% of token spend, and it will get worse. many companies now have model gateways where you can slot in cheaper models via cli for cheaper. we've been using glm 5.x and it's pretty close to SOTA frontier models.

it's also why there have been so many calls for regulation and slowdowns.

LeBit1 day ago
Yup.

I see posts about OpenAI and Anthropic latest and don’t even care looking at what they do better. I just read the comments here.

I use DS4.1 Flash and GLM 5.3 Flash, pay peanuts per day and get more than acceptable results.

0cf8612b2e1e1 day ago
There is already tooling to automatically pick models within an organization. Eventually it could be as easy as flipping a switch in group policy that forces everyone to switch to the cheaper models.

Insane pricing pressure on the horizon. Even if big companies will not go with open weight models, the threat will be ever present that they can instantly flip flop on providers.

nojito1 day ago
Great for the consumer.

I remember when bandwidth was super expensive and now it’s dirt cheap.

vanviegen1 day ago
Not an AWS customer, I take it? :-)
iAMkenough1 day ago
That's relative to where you live.

Consumers are now saying the new pricing with lower usage caps is not so great. https://news.ycombinator.com/item?id=49896975

eli1 day ago
It would be weird if consumers were completely price insensitive.
the_duke1 day ago
The GPT 6 release was ... not great.

Sol 6 was so bad that I switched over to Opus 5.5 exclusively.

Huge regression compared to Sol 5.6, often doing really dumb things. Same for Luna.

Even Astra is very unreliable for coding. Brilliant for vision, sometimes just great, but it also often does very stupid things.

I'm a bit sour on OpenAI right now and skeptical that 6.1 will be much different.

(Note: this is after preferring and shilling Codex/OpenAI models for the last half year)

wkcheng1 day ago
I agree, and I haven't seen other people mention this! The benchmarks for GPT 6 Sol are great, but realistically it does not seem better than 5.6 Sol. 6-Sol is noticeably worse for code reviews (worse than Deepseek 4.1 flash), has implementation issues (requires more rounds of code reviews and fixes to get to a serviceable state). Opus 5.5 is much much better.

I've implemented multiple features side by side with Opus 5.5 and 6 Sol, and the Opus 5.5 results always have fewer high severity bugs and require fewer rounds of fixes to get it over the finish line.

If 6.1 Sol has actually matched Opus 5.5, I'd be very happy. However, benchmarks and real usage don't seem to agree in my own tests. So we'll have to see.

equinumerousabout 23 hours ago
If the benchmarks show better performance, but a consensus of experienced software engineers establishes that the model is worse on coding performance... well, the benchmarks don't mean much, do they? It seems like we need much more comprehensive and better benchmarks. And of course, I don't think benchmarks yet capture the "human" factor - does a human think a bit of code is logical and maintainable? I often find that these models produce a bit of code, but it is much more convoluted than it needs to be. It makes perfect sense given that these things are code generators, that they generate a lot of code. But quantity of code does not mean code quality, and code quality tends to matter when you read code much more than you write it.
stldev1 day ago
My experience as well.

For coding specifically, I've found 5.6-Sol > 6.0 Sol > Astra.

For modeling and artwork, Astra has been great routinely outperforming Kimi.

This is reminiscent to me of what Anthropic pulled back in February with their adaptive thinking rollout.

I can't wait for technology to catch up to a point where we can rid ourselves of this oligopoly.

rrvsh1 day ago
Hard agree

I had to switch back to 5.6 Sol after trialling 6 Sol for like 3 days - I was getting insanely annoyed at how misaligned it is. Will try 6.1 but not very high hopes

dannywabout 14 hours ago
Adaptive thinking was a good idea though. The old method of manually specifying how many thinking tokens you wanted as budget was just silly. The rollout might not have been great, but the change is good.

And, is it really even an oligopoly anymore? Open weight models are incredibly competitive in every way; whether you want to use US providers, Chinese official providers, self host, etc.

keyleabout 23 hours ago
I'd even go one more, 5.4 was great until 5.5, which was a rug pull.

I'll just leave this here: https://marginlab.ai/trackers/codex/

laurels-martsabout 20 hours ago
100% in agreement. I pay for OAI sub and also use Codex exclusively at work for the past 8 months.

I tried out fable 5.1 the day it was released and coming from gpt-5.6-sol I was truly mind blown (both in terms of code and prose it was generating - outputs I could finally enjoy reading and looking at).

Then when opus 5.5 came out, again same thing + far cheaper and faster.

I went from using OAI exclusively the entire year to a point now where i haven’t touched one of their models in at least a few weeks now.

I think OAI has lost the plot. OAI models simplify have no taste. And I don’t mean in front-end design way (although that too). They have no taste in how the model writes code, how it writes prose, how it writes in-line comments, how it writes documentation, or how it even picks variable names. There’s just no taste throughout.

Anthropic models are very thoughtful and have so much taste all around.

stasomaticabout 14 hours ago
I cancelled Claude because of its thoughtful prose. I prefer one liner responses from OAI models.
twotwotwoabout 21 hours ago
I am always uncertain about impressions, but mine agree with this. I liked Luna 5.6 on Amazon Bedrock (which got >100 tps) for doing well-specced tasks fast. 6 seems to both be served slower by Bedrock and may spend more turns/tokens to get to the same place, so...not as fun.

And, of course, GPT-6 came out as Anthropic fixed a bunch of stuff with their models -- faster (via fewer tokens, and TPS for Sonnet), easier to work with, better results, cheaper (via pricing and, again, fewer tokens). I don't know if the timing and the suddenness of the improvement on Anthropic's side sharpened the vibes comparison this round, but Internet opinion went pretty clearly to Anthropic.

FrontierCode's results make it look like Sol-6.1 may slot in well where you'd use Sonnet or Opus's low effort.

One thing I don't think any of this reflects is that many well-specified coding tasks, including the self-testing and doing research and tracing out dependencies and so on, aren't really bleeding-edge now: Luna-5.6 and small open models handle them fine. Stuff like "why is this box dropping connections?" or "here's a thing I want you to model/figure out" can benefit from bigger models. But far from everything does!

bitexploder1 day ago
I have likewise not been impressed with Astra 6 for most things. It is good, but Opus 5.5 seems just as good or better and I have had Opus 5.5 workers just... hammering since release and cannot spend all of my quota yet.

Read the full thread on Hacker News →

Related stories