GPT-6.1 Sol replaces GPT-6 Sol after just 7 days. It scores 1 point below GPT-6 Astra in the Intelligence Index at less than one quarter of the Cost per Task

67 points•theanonymousone•about 15 hours ago•84 comments•

84 comments

atkristaabout 14 hours ago
I would just LOVE to see all the behind-the-scenes shithousery both companies are employing to one-up the other in this, largely, 2-horse AGI race. Someone should make a mockumentary when all is said and done!
nbardyabout 14 hours ago
I think it's weirdly just a choice of deciding to cut releases.

We already know OpenAI has "bel" that is MUCH better than astra and is being used internally

233mhzabout 13 hours ago
> We already know OpenAI has "bel" that is MUCH better than astra and is being used internally

I mean, isn't it almost a guarantee that what we get is a gimped version of what they use internally? They probably already serve themselves next gen level models at 1k+ tps from cerebras machines hosted on perm while we get quantized astra/opus at 50tps on a good day

iLoveOncallabout 14 hours ago
> We already know OpenAI has "bel" that is MUCH better than astra and is being used internally

You're just believing their own bullshit. There's no indication that this is true except from claims from people working at OpenAI.

If they really had a much more powerful model, it would make absolutely no sense to sit on it.

mFixmanabout 14 hours ago
Any strong enough model with weak enough safeguards can cause an AI Chernobyl event that will make people and governments against AI development and deployment, just like Chernobyl did for nuclear energy.
petesergeantabout 14 hours ago
> largely 2-horse

The absolute frontier is largely 2-horse, but the rest of the pack is very close behind, which I'm grateful for. Grok, Facebook, and the Chinese vendors are producing excellent models.

bayindirhabout 14 hours ago
Gemini is also pretty nice for researching things. It turns out that having the whole internet indexed and having unlimited access to YouTube is a force multiplier of some kind.

Since Google has their own TPUs, TPS is also pretty high w.r.t. Claude, for example.

Bluesteinabout 14 hours ago
... and, must be said a plethora of largely unsung, small, unknown "labs", outfits, "researchers" and the like. There is a long tail of smart people having at this. I guess, sheer compute aside, I think much progress - or, at least, important pieces thereof, will come from there.-
michelbabout 14 hours ago
I REALLY need new seasons of 'Silicon Valley'
tom1337about 13 hours ago
Still waiting for the moment where GPT either orders 4,000 pounds of raw beef or just deletes the whole OpenAI repo because it "thought the easiest way to remove all bugs is by deleting the whole repository"
TeMPOraLabout 14 hours ago
AGI will make one, about humanity, after we're all gone - "They were so dumb, they just deserved to die".
f6vabout 14 hours ago
The sooner the better, brother.
bearjawsabout 14 hours ago
What do people get out of being so reductionist?

Every time a new model comes out, people come out in droves "oh I don't notice anything different".

People have been saying this about <currentModel-1> for 2 years now, and the entire state of AI has changed dramatically.

It cannot be that the next AI model isn't better, but also suddenly what they are capable of is on an entirely different level.

ameliusabout 14 hours ago
I suppose it's the reverse of Amara's law:

"We tend to overestimate the effect of a technology in the short run and underestimate the effect in the long run"

ulimnabout 14 hours ago
I suspect it's partly because people didn't jump from GPT-3 to GPT-6.1 Sol and partly because SOTA models from the last few(?) months have been able to tackle most of the regular tasks. It means this new model isn't different in that regard from Opus 4.8, if your mental benchmark is that they both are capable of implementing something like a CRUD app.
alstoniteabout 14 hours ago
I’ve seen overwhelmingly that when a model is good people see it, and when it isn’t, they criticize. I’ve seen nothing but ‘wow this is a huge step up’ from Opus 5.5. I felt this way about Opus 4.5, GPT-5.6 Sol/Luna, and to a lesser extent with Fable and Astra. But Opus 5 was ass, and the entire gpt 6 line feels like OpenAI’s version of that.
singpolyma3about 13 hours ago
They've been pretty capable for a very long time. I don't think the models are getting more capable so much as people are getting better at using them and more people are getting the opportunity to be impressed.
sosukeabout 14 hours ago
I like the fast releases. 6.1 coming so fast after 6.0 means they found some improvement solid enough for a new rollout.

The only time I remember the newest model making an obvious regression was when the first rolled out MoE. Super speed update but each request had less intelligence at hand. We’re way past that now

hgoelabout 14 hours ago
It seems strange though.

Did they find this improvement within the week? If so, given all their whinging about safety, it seems irresponsible to only test the improved model for less than a week.

Did they find the improvement more than a week ago? If so, why bother releasing GPT 6 if they knew they had a better version essentially ready to go?

roody15about 11 hours ago
I suspect google has a model that is at the same level but is waiting strategically to release rather than this constant week to week dash.

Still seems to me like google is in a good position with AI just based on their size and approach.

BugsJustFindMeabout 6 hours ago
Why speculate?
egeozcanabout 15 hours ago
Worries of AI going rogue take so much attention that no governance body seems to care about the shady subscriptions and limits business.
entropeabout 14 hours ago
What about those do you think are "shady"? Price discrimination in favor of small customers at the expense of large customers is somewhat common; businesses want customers to buy more of their products, but customers are not obliged to buy more if they like their current deals. Having limits on using finite resources seems even easier to justify.
egeozcanabout 14 hours ago
Shady as in not clearly communicated and change with random tweets being the only indicator, or a blog post if you're lucky.

x20? 25x? 5x? All in comparison to some other plan that also doesn't have clear limits. "Oh BTW we mean session limits, your weekly is 2x of 5x". "This model eats your limits twice as fast and you can use half of your weekly limits on it". "No this time we mean weekly only".

SkyBelowabout 14 hours ago
Humans see differences in prices as unfair. The greatest example of this is price gouging during an emergency, but also look at the level of hate that scalpers get.

Using AI to do this, if anything, given the common negative sentiment, is seen as even worse. People think of this as AI using information asymetry to squeeze more out of users, not to cut people a deal.

Taking advantage of information asymmetry is generally looked down upon as well. Look at all the laws we have protecting kids from this. Businesses often don't seem similar protections because those are businesses with big legal teams (and when it is a big legal team vs a small mom and pop store without a single lawyer on payroll, people do start taking issues with it). The power difference between the average company using AI pricing and the average consumer falls pretty solidly in the 'we don't accept this' side of taking advantage of information asymmetry.

I could keep going, but I think these are already plenty enough reasons to why people look at AI price discrimination as not just a bad business practice they don't like, but an immoral/unethical one.

Sure, one can make economical counter arguments, but that's arguing on an orthogonal dimension that simply isn't relevant to where these feelings/thoughts come from.

semiquaverabout 14 hours ago
Do you mean the subsidized subscriptions which allow individuals to pay a tenth of the normal API cost for tokens?
muyuuabout 14 hours ago
Dumping is shady af.
Dinudaabout 15 hours ago
After 5.5, I basically don't notice a jump in model performance, other than my usage ending sooner.
redox99about 12 hours ago
Because you're probably using it for stuff like "edit this function"/"refactor this class". 5.5 to 5.6 Sol was a giant jump. 5.6 to 6.1 seems very large as well.
user43928about 14 hours ago
I do.

The results are less buggy, animations are much better.

It can work autonomously for hours and the result is decent most of the time.

That wasn't usually the case with 5.5, which needed more feedback and iterations to get things right.

baqabout 14 hours ago
Astra is noticeably smarter than any OpenAI model before it. Sol 6.1 is very noticeably smarter than sol 6 even after half a day of using it (sol 6 was actually terra 6 and opus 5.5 has taken them by a total complete surprise)
rimliuabout 14 hours ago
none of them is smart.
tom1337about 15 hours ago
Kinda same but I miss my 5.3 Codex. Thing lasted forever on my $20 subscription and with detailed prompts was able to pretty much implement everything I requested it to do with a acceptable quality.
zero1009about 15 hours ago
Felt the same until I started using Luna. I feel like I get similar performance, but faster, and my usage lasts so much longer.

Read the full thread on Hacker News →

Related stories