Claude Sonnet 5.5 is a clear upgrade over Claude Sonnet 5, runs 30%+ faster, and costs up to 30% less for most work.

880 points•D2OQZG8l5BI1S06•2 days ago•609 comments•

609 comments

Sol-2 days ago
Probably a first world problem, but with Opus 5.5's efficiency, the limits on the 5x plan are simply sufficient for my everyday work, even when running 2-3 sessions at a time. So I wonder when I would use Sonnet 5.5.

More concurrency than that isn't really practical for me if I want to retain some semblance of understanding. Perhaps it's different for purely web app or frontend tasks, where the outcome is more relevant than the process, I don't have much experience there (and also don't want to belittle these domains, I might be underestimating their complexity).

So surprisingly, my own work is at least for the time being almost saturated by the model capabilities. I am not sure how I'd scale from here. Sure I could run all requests at max effort to burn tokens for the sake of it, but that can't be it. And for many tasks, I am not really able to define so clear cut success criteria or self-verification loops that I could benefit from letting an agent (or a fleet thereof) autonomously run for a day.

So I realize it's a skill issue on my side, but I can't be the only one. I wonder if there is a limit to token demand, at least short term. Feels like either they accelerate to AGI and RSI, where the AI can find uses for token, or things might plateau at some point.

Note I don't think this because I'm an AGI skeptic or think there's a ceiling to intelligence, but there might simply be a valley of economic hardship for the companies where the supply of tokens outpaces the demand, due to a lack of ideas of what to do with them. And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.

miki1232112 days ago
I find that "vibe coders" (that is, people who do not know anything about programming, but nevertheless produce useful tools for themselves and others) are using a lot more tokens than we do as programmers.

I think this is partially because we're still attached to pre-LLM notions of architecture, good design and code quality (which are still important, but maybe less important than they once were and that we think they are), partially because their projects are in a messy state, so models have to work around the technical dept.

They're essentially trading off programmer time for LLM time (which is a good trade financially speaking).

inopinatus2 days ago
It's because they don't know data structures.

"Show me your flowcharts and conceal your tables, and I shall continue to be mystified. Show me your tables, and I won’t usually need your flowcharts; they’ll be obvious." - Fred Brooks, The Mythical Man-Month (1975).

and essentially the same sentiment, three decades later:

"Bad programmers worry about the code. Good programmers worry about data structures and their relationships." - Linus Torvalds, git mailing list, 2006.

These things have not changed even though everything else is topsy-turvy. As-of current writing, I have yet to see an LLM make good data structure choices; they go for something that is superficially plausible but profoundly ill-considered (or rather, not considered at all), and then commonly burn tokens treating this implementation detail as a design invariant and trying to deal with the consequences by writing more code, instead of iterating directly upon the ill-fitting data at the root its problems.

If you're wondering, "does he mean the schema of let's say a db or other persistent store, or does he mean abstract/algebraic structures", the answer is yes to both, I think coding models are today shockingly weak when it comes to design reasoning in both domains.

Fortunately, their suggestibility means the same models will readily accept direction on the matter (perhaps even more so than on the structure of code), so I recommend doing just that, and (bonus!) this means your CS degree is still relevant.

gchamonlive2 days ago
I think it's not only a matter of token efficiency. If you don't know what you are doing development will eventually crawl to a halt invariably.

It's the compound counter-probability of success, so even a 99% efficient model will in time accumulate so much error that without conscious cleanup and steering, it becomes really unlikely really fast that anything could be changed in the code without affecting something else, no matter how many tokens you throw at it. It's the collapse of a complex system under the weight of sheer uncertainty of what the system actually does.

gobdovan2 days ago
I think this is valid now, but not guaranteed to be valid forever. For engineers, there was a period where more checks, more tests, more auto code reviews improved results quite a bit. People were consuming tokens like crazy (including me). Then things improved via better effort/thinking levels, where you could see repeated code reviews plateaued, so now people don't really do that quite as much.

There was also a period where specifically OpenAI models would always have to comment something in code review and the builders were agreeable up to listening to each nitpick. If you'd have a loop of build->review->build->review, it would take maybe 5-7 rounds for it to 'settle' and not find the smallest nitpicks to argue about. Tried it this week with Astra reviewer and it's about 0-2 review loops (never had a LLM accept a change without nitpicking first try before Astra).

There was also a period where you'd have to give quite specific instructions for agents to keep iterating, but now agent are pretty proactive and try to finish tasks you give them unsurprisingly most of the time.

So, while there's a shortcoming of LLM+harness and engineers observe more tokens improve things even logarithmicly, you'll see more tokens seemingly abused by engineers.

furyofantares2 days ago
For my normal work I take ownership of the code, and end up with the exact code I want. I still have it go off and do a good amount of work a lot of the time, still queue up multiple tasks at the same time a lot of the time. Sometimes I throw it all away and re-prompt once it's time to commit to it, sometimes edit what it made, sometimes have it edit what it made etc.

For all of my side projects I'm full-on vibe. Well, almost: I do have opinions on what kinds of code it should write and set up my projects to get that. But I don't LOOK at the code.

I use a LOT more tokens on my side projects. I can have it working more or less constantly and it doesn't take up that much of my attention, but it is FAR less token efficient.

arceister1 day ago
Because that "vibe coders" didn't know and go through the fundamentals, thus they're wasting tokens with probably continuing the AI hallucination suggestions.

I've seen bunch of persons like this and that's kinda stupid because they're just blindly following AI's "suggestions" while they actually don't know what they're doing, then results on terrible code and architecture with "if it works, it works" mentality.

maherbeg2 days ago
There's lots more you can do! Use the model to monitor your deployments after they get deployed. Have them fix and watch CI issues for you. Run adverserial review. Automatically watch metrics every day and highlight performance regressions. Start reviewing your previous sessions to find ways to statically reject different failure modes and have the agent have more success earlier on etc.

Another thing to think about is, what would it take for you to care less about the understanding. Better integration / e2e tests? Performance validation? visualizing program and data flows? Better refactoring of your modules?

mattm2 days ago
> what would it take for you to care less about the understanding

It's an interesting question. The thing I keep coming back to though is that every time I've tried to go more towards vibe-coding, I invariably look at the code and find things have been added that would just not be acceptable. I've also tried asking the models to see could be refactored however they still miss things that should be obvious.

I think the gap is that they're still lacking a sense of importance. As engineers working on a product, you have a sense that this feature is more important than that feature. An LLM treats your codebase at the same level of importance. So they'll spend the same amount of effort and code changes on testing and hardening something that just really isn't that important.

Also, once a bad pattern gets into the codebase, they just continue to build and extend that out rather than re-thinking about it like an engineer would.

klardotsh2 days ago
The thing with watching CI in an agent loop is that it burns tons of tokens. At work I ended up writing a deterministic, traditional CLI tool to poll GitLab CI pipeline+job state changes on a branch and exit with an appropriate status code, and then updated my `/glab-ci-feedback` skill to use that. Saved a ton of token churn, and now I have a runbook a human could just as easily use if they don’t want to (or can’t) use an agent loop.

… but walking away to make a coffee and coming back to the robots auto-fixing bugs only found in CI is definitely some flavor of magic, regardless of the execution order to get there.

miki1232112 days ago
I think that's what a future dev team is going to look like.

One person doing product management / talking to customers and vibe coding features that solve users' problems, one person keeping the UI/UX in check, one QA person that spends their time clicking through the software, finds the bugs that are obvious to humans but not LLMs and fixes them, and one "harness engineer" who pays off technical debt, observes failure modes and sets the rest of the team up for success.

Hauthorn2 days ago
> Another thing to think about is, what would it take for you to care less about the understanding.

Could you explain why it would be a goal to understand the system less, rather than more?

It seems harder to know if you have good tests while lowering your expertise in the system.

tshaddox2 days ago
More tests that aren’t written by you don’t help you understand the system, and I would argue the there’s no confidence without understanding. That was true in the pre-agentic era and is perhaps even more true now.
phainopepla22 days ago
It's the "semblance of understanding" you're holding on to that is keeping your demand limited. I'm holding onto it as well, but I think these companies are assuming that human understanding will no longer be relevant for most codebases going forward.
Imustaskforhelp2 days ago
> It's the "semblance of understanding" you're holding on to that is keeping your demand limited. I'm holding onto it as well, but I think these companies are assuming that human understanding will no longer be relevant for most codebases going forward.

In short, seems to describe vibe-coding to me? What I don't understand about companies attempting to vibe code is if they realize that other people (especially sometimes their customers) can tailor-made their own software for their own needs, or rather competitors can be dime a dozen and maybe even a fight for constantly paying for the better model.

There was a comment[0] from a just few days ago by @jjcm (which I wish to quote which I hope they don't mind.):

> I just got back from a 2 week trip to China. I was in some of the more remote parts and my cell wasn't able to connect to their towers in that area, resulting in me not having the tourist VPN.

> The side effect was I was fully cut off from my AI tools for those two weeks. I was coding "manually" during that time, and I think I accompished in two weeks what I previously had been able to do in a day. I'm not gonna lie, it was very, very stressful as a solo founder.

> The industry moves so fast these days, that the only way to keep up with the speed is to leverage them. While I can appreciate the push of this to help your brain think independently/critically, the opportunity cost of a month of development without LLMs is too high a price to pay.

What happens if the opportunity cost of a month of development with vs without human understanding becomes too high a price to pay. I feel like we would be in awkward time because of the factors that I had described above (higher competition, software stops meaning just as much software as people would be custom-making them.)

I think that (former fly.io's) @tptacek's article[1] starts making more sense if viewed from this direction: What even is an OS now.

I don't have the answer to this question as to what happens next but its a form of development that I would prefer not to happen on a more gut instinct level?

Letting AI basically control everything and us not having any mental understanding of sorts and sort of becoming the meat-proxies just for economical reasons seems realistic possibility but a bleaker reality at that. I am left feeling a little bit uncomfortable if this reality turns out to be true.

[0]: https://news.ycombinator.com/item?id=49808422

[1]: https://sockpuppet.org/blog/2026/09/25/what-even-is-an-os-no...

andrepd2 days ago
Damn, yet they still hire programmers, marketers, researchers like there's no tomorrow. I thought everything would be vibe coded and we wouldn't need to even understand code anymore. Which one is it?

The proof of the pudding.

egeozcan2 days ago
I created a team of agents using Opus 5.5 to review and address findings on a job system I have in a side project with medium reasoning, and I burned through the 20x plan weekly limit in 2.5 days. They were using GPT-6-Sol for reviews, and it also used 85% of my OpenAI x5 weekly limit. Three hundred something commits in total.

OTOH, in the daily job, I have the team plan that's similar to 5x plan and I never had any limit problems, because I really need to understand be able to take responsibility for the code.

Totally different uses.

gregwebs2 days ago
> I want to retain some semblance of understanding

How you do this (and how deeply) I think is really the limit. I am doing this by focusing heavily on the design phase with grilling and trying to continually improve process to need less effort in the review phase. Are your models doing automated reviewing and testing before pushing out the PR (themselves)?

I think in the long run as models and the tools around them get better and cheaper, those that abdicate understanding will be able to achieve more. Although programmers think of that as irresponsible, ask yourself what does a tech lead do? And then what does a CTO do, etc?

jwpapi2 days ago
I think going for more understanding is the way you need less understanding. The more solid your core understanding of your codebase is the less you need to know the details, the less missunderstandings the less iterations needed, the less mental capacity consumed
thefourthchime2 days ago
It does vey well at one shotting a PacMan clone, pretty much perfect. https://jonclegg.github.io/pacman-bakeoff/entries/claude-son...

2nd only to Opus 5.5, which is perfect. https://jonclegg.github.io/pacman-bakeoff/entries/claude-opu...

Up until very recently, all models struggled with this.

All results: https://jonclegg.github.io/pacman-bakeoff/

judge20202 days ago
Oh, it coded a Pac-Man clone. The clone was so good that I thought it was premade in some way and that Sonnet was going to play PacMan.
thefourthchime2 days ago
Yes! The point being that up until yesterday, every model struggled with this, and now they don't.
sixtyj2 days ago
I have played few of them and it seems that Opus 5.5 is the first one who really made playable PacMan clone game. On mobile as well.

Could it be because the model was somehow pre-trained? If we compare it with pelicans that are still not-perfect…

lukan1 day ago
" If we compare it with pelicans that are still not-perfect…"

Fable 5.1 was pretty good. Even animating it:

https://news.ycombinator.com/item?id=49526704

Well, I thought you meant the Pacman package manager for a second, still impressive. I've noticed the last few versions of Opus have produced better video game coding output.

https://wiki.archlinux.org/title/Pacman

russellbeattie2 days ago
Wow, that "bake off" page is better than any coding benchmark I've seen! You can really sense the strengths and weaknesses of each model/harness combo.
thefourthchime2 days ago
Thanks!
nicce2 days ago
I was able to get similar with Qwen 3.8 27B with one shot. I think this game is too well in the training data.
abejora2 days ago
Sonnet 5.5 scoring higher (70.6) than Opus 5.5 (66.4) in Terminal-Bench is interesting. I looked into this, because it felt strange.

Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.

[1] Section 8.5 of the Sonnet 5.5 System Card

eli2 days ago
Why isn't that worth reading into? I care about the experience of actually using the model, not hypothetically what it could achieve without overactive guardrails
abejora2 days ago
You're right about its real world performance, and I worded my original comment wrongly.

I was merely thinking of the theoretical aspect of it: performance of opus 5.5 is better than sonnet 5.5 across the board, with the exception of Terminal-Bench. So I was curious why this one stood out. Was it because they focused on it during training? Did sonnet 5.5 had access to more references for this benchmark? But based on my first reading, I concluded that it might just be the safety constraints that made the difference here, and I wanted to share that.

spider-mario1 day ago
You can read “a little bit” (i.e. “not too much”) into it (it does indeed tell you about the out-of-the-box experience), but e.g. being able to know when a fallback model has been used means that in terms of pure accuracy, you might still be better off defaulting to Opus 5.5 and re-routing to Sonnet 5.5 yourself when you get the fallback.
chis2 days ago
Well presumably now it’ll fall back to Sonnet 5.5 lol
subscribed2 days ago
I disagree, I think we should read a lot from it, as it stands in this benchmark Opus performs worse than Sonnet, it doesn't really matter why.

Anthropic made it that way, and I'd say the lower score is accurate.

bitexploder2 days ago
If you care about the things terminal bench cares about, yes. Sonnet was probably trained aggressively on agentic coding and things that align well with deepswe and terminal bench and or tuned heavily for those tasks. Sonnet is an agent likely to do more of those tasks and be given the more grunt work tasks. Whilst Opus' wider knowledge pool means it can deal with a much higher variety of real world situations successfully. And, those benches are often timed or limited. Opus may have been running out of time. Looots of factors.
cromka2 days ago
It matters if it's not Sonnet performing the task, doesn't it?
Leary2 days ago
And Sonnet 5.5 is more expensive than Opus 5.5 to hit that score on terminal bench!
radlad2 days ago
I believe you meant to cite the Opus 5.5 System Card which states:

> Claude Opus 5.5 scored 66.36% on Terminal-Bench 4.0 with safeguards enabled; requests flagged by the safeguards were answered by a fallback model following the default server-side fallback policy (2.5% of requests, affecting 10% of trials).

> https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba50242199...

I cannot find a Sonnet 5.5 system card.

oh_no2 days ago
it could be that, it could also be that sonnet max looks to burn about 60% more tokens than opus max

AA intelegence index (agent harness doesn't have sonnet data yet) on max: Astra 27k Fable 5.1 78k (Sonnet 5) 118k Opus 5.5 119k Sonnet 5.5 193k

Opus 5 was previous record holder so hats off to Anthropic on blowing it away on token churn.

azuanrb2 days ago
Unless you’re using frontier models like Astra, Sol, Fable, or Opus, I think you’re often better off using Chinese models for a fraction of the price. I’m not sure people realises just how competitive they’ve become.

GLM and DeepSeek are great examples. They’re a bit like Linux or Android in that there isn’t necessarily one best provider. You need to do some research, try a few, and pick whatever works best for your use case.

I think that’s partly why Anthropic has been pushing its most expensive models so heavily for a while now. Sonnet and Haiku were great, but at that level of intelligence it’s becoming much harder for them to compete on price with Chinese models that have largely caught up.

The main reason to use frontier models from Anthropic or OpenAI now is the combination of intelligence and speed. Chinese frontier models still struggle to match that, possibly in part because of hardware constraints. But judging by the recent GLM releases, they seem to be moving in the right direction.

persedes2 days ago
Those models are cheaper per token, but depending on your use case you might still end up paying more with the cheaper model. DeepSeek likes to burn through a lot of tokens for example, which can quickly ameliorate those savings. They are great models and have most likely helped anthropic and openai drop their prices lately, but I still don't see the monetary benefit of using them atm.
noisy_boy2 days ago
That has been my experience with DeepSeek after the price increase. I was so used to it being so frugal, it took couple of top-ups for me to realize that it wasn't the budget-king anymore.
grahamnorton392 days ago
You might mean “eliminate” or similar - to ameliorate something is to improve it or better it :)
benced2 days ago
Luna is the main exception to this. It's such a cheap model and the tokens come so fast.
tripleee2 days ago
Luna is genuinely my favorite model. I like developing in small chunks instead of huge sweeping changes, and Luna is so good for that.
pkulak2 days ago
I cancelled my codex subscription 2 days ago, but only because I knew I could wire into Luna API pricing. Luna xhigh writes very good code, basically for free.
shepherdjerred2 days ago
Yup. Luna is fast, cheap, and pretty intelligent. IMO the Codex harness is a bigger limiter than the model itself
scuppernong2 days ago
if you're paying API prices, you probably already know this. everyone else is using a subscription which is massively subsidized rel API prices. or am I missing a third case?
phoghed2 days ago
Does anyone beat Luna on price? It’s surprisingly capable for a lot of things.

Most enterprise customers are paying per token at this point afaik, whether that’s to gh copilot, Anthropic, or running models on Vertex/Azure/Whatever

madeofpalk2 days ago
Third case is you’re just an employee at a company who pays for an enterprise codex/claude/whatever for you, and you just use whatever the best available is. Maybe they’ve locked away Astra or Fable, but if cost literally isn’t an issue (at this moment) is there still a benefit to Deepseek?
raincole1 day ago
Chinese models are not for a fraction of the price though. If you only have a budget of $20 a month, luna with chatgpt subscription actually gives you much more room than deepseek. The idea that Chinese models = cost efficiency is rather outdated.
AbstractH241 day ago
It's still unclear to me how much I'd have to use a Chinese model to equal my Claude Max subscription price.

There's something to say for flat pricing rather than per token. Even if its not a better deal.

vincengomes2 days ago
For the folks asking what is the point of Sonnet 5.5 when Opus 5.5 is better in every way, Sonnet is now the new default in the free tier and now people using Claude Web in the free tier get access to an almost close to the frontier model.
ddxv1 day ago
I've been using the Claude free tier (browser chat) for awhile and never seem to hit restrictions anymore. There used to be many more restrictions a year ago. At the same time, if they enforce a minimum paid tier I'd probably just switch to Gemini/ChatGPT or whatever other free model is around at the time.
etatester1 day ago
Hello fellow copy-paster. I hit them regularly when asking for medium refactors with higher effort. I have ChatGPT in another tab for when that happens, but to me it's awful for anything non-trivial. Gemini I don't even consider unless I'm just asking "what would Gemini do"

Read the full thread on Hacker News →

Related stories