Claude Sonnet 5.5 is a clear upgrade over Claude Sonnet 5, runs 30%+ faster, and costs up to 30% less for most work.
609 comments
More concurrency than that isn't really practical for me if I want to retain some semblance of understanding. Perhaps it's different for purely web app or frontend tasks, where the outcome is more relevant than the process, I don't have much experience there (and also don't want to belittle these domains, I might be underestimating their complexity).
So surprisingly, my own work is at least for the time being almost saturated by the model capabilities. I am not sure how I'd scale from here. Sure I could run all requests at max effort to burn tokens for the sake of it, but that can't be it. And for many tasks, I am not really able to define so clear cut success criteria or self-verification loops that I could benefit from letting an agent (or a fleet thereof) autonomously run for a day.
So I realize it's a skill issue on my side, but I can't be the only one. I wonder if there is a limit to token demand, at least short term. Feels like either they accelerate to AGI and RSI, where the AI can find uses for token, or things might plateau at some point.
Note I don't think this because I'm an AGI skeptic or think there's a ceiling to intelligence, but there might simply be a valley of economic hardship for the companies where the supply of tokens outpaces the demand, due to a lack of ideas of what to do with them. And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.
I think this is partially because we're still attached to pre-LLM notions of architecture, good design and code quality (which are still important, but maybe less important than they once were and that we think they are), partially because their projects are in a messy state, so models have to work around the technical dept.
They're essentially trading off programmer time for LLM time (which is a good trade financially speaking).
"Show me your flowcharts and conceal your tables, and I shall continue to be mystified. Show me your tables, and I won’t usually need your flowcharts; they’ll be obvious." - Fred Brooks, The Mythical Man-Month (1975).
and essentially the same sentiment, three decades later:
"Bad programmers worry about the code. Good programmers worry about data structures and their relationships." - Linus Torvalds, git mailing list, 2006.
These things have not changed even though everything else is topsy-turvy. As-of current writing, I have yet to see an LLM make good data structure choices; they go for something that is superficially plausible but profoundly ill-considered (or rather, not considered at all), and then commonly burn tokens treating this implementation detail as a design invariant and trying to deal with the consequences by writing more code, instead of iterating directly upon the ill-fitting data at the root its problems.
If you're wondering, "does he mean the schema of let's say a db or other persistent store, or does he mean abstract/algebraic structures", the answer is yes to both, I think coding models are today shockingly weak when it comes to design reasoning in both domains.
Fortunately, their suggestibility means the same models will readily accept direction on the matter (perhaps even more so than on the structure of code), so I recommend doing just that, and (bonus!) this means your CS degree is still relevant.
It's the compound counter-probability of success, so even a 99% efficient model will in time accumulate so much error that without conscious cleanup and steering, it becomes really unlikely really fast that anything could be changed in the code without affecting something else, no matter how many tokens you throw at it. It's the collapse of a complex system under the weight of sheer uncertainty of what the system actually does.
There was also a period where specifically OpenAI models would always have to comment something in code review and the builders were agreeable up to listening to each nitpick. If you'd have a loop of build->review->build->review, it would take maybe 5-7 rounds for it to 'settle' and not find the smallest nitpicks to argue about. Tried it this week with Astra reviewer and it's about 0-2 review loops (never had a LLM accept a change without nitpicking first try before Astra).
There was also a period where you'd have to give quite specific instructions for agents to keep iterating, but now agent are pretty proactive and try to finish tasks you give them unsurprisingly most of the time.
So, while there's a shortcoming of LLM+harness and engineers observe more tokens improve things even logarithmicly, you'll see more tokens seemingly abused by engineers.
For all of my side projects I'm full-on vibe. Well, almost: I do have opinions on what kinds of code it should write and set up my projects to get that. But I don't LOOK at the code.
I use a LOT more tokens on my side projects. I can have it working more or less constantly and it doesn't take up that much of my attention, but it is FAR less token efficient.
I've seen bunch of persons like this and that's kinda stupid because they're just blindly following AI's "suggestions" while they actually don't know what they're doing, then results on terrible code and architecture with "if it works, it works" mentality.
Another thing to think about is, what would it take for you to care less about the understanding. Better integration / e2e tests? Performance validation? visualizing program and data flows? Better refactoring of your modules?
It's an interesting question. The thing I keep coming back to though is that every time I've tried to go more towards vibe-coding, I invariably look at the code and find things have been added that would just not be acceptable. I've also tried asking the models to see could be refactored however they still miss things that should be obvious.
I think the gap is that they're still lacking a sense of importance. As engineers working on a product, you have a sense that this feature is more important than that feature. An LLM treats your codebase at the same level of importance. So they'll spend the same amount of effort and code changes on testing and hardening something that just really isn't that important.
Also, once a bad pattern gets into the codebase, they just continue to build and extend that out rather than re-thinking about it like an engineer would.
… but walking away to make a coffee and coming back to the robots auto-fixing bugs only found in CI is definitely some flavor of magic, regardless of the execution order to get there.
One person doing product management / talking to customers and vibe coding features that solve users' problems, one person keeping the UI/UX in check, one QA person that spends their time clicking through the software, finds the bugs that are obvious to humans but not LLMs and fixes them, and one "harness engineer" who pays off technical debt, observes failure modes and sets the rest of the team up for success.
Could you explain why it would be a goal to understand the system less, rather than more?
It seems harder to know if you have good tests while lowering your expertise in the system.
In short, seems to describe vibe-coding to me? What I don't understand about companies attempting to vibe code is if they realize that other people (especially sometimes their customers) can tailor-made their own software for their own needs, or rather competitors can be dime a dozen and maybe even a fight for constantly paying for the better model.
There was a comment[0] from a just few days ago by @jjcm (which I wish to quote which I hope they don't mind.):
> I just got back from a 2 week trip to China. I was in some of the more remote parts and my cell wasn't able to connect to their towers in that area, resulting in me not having the tourist VPN.
> The side effect was I was fully cut off from my AI tools for those two weeks. I was coding "manually" during that time, and I think I accompished in two weeks what I previously had been able to do in a day. I'm not gonna lie, it was very, very stressful as a solo founder.
> The industry moves so fast these days, that the only way to keep up with the speed is to leverage them. While I can appreciate the push of this to help your brain think independently/critically, the opportunity cost of a month of development without LLMs is too high a price to pay.
What happens if the opportunity cost of a month of development with vs without human understanding becomes too high a price to pay. I feel like we would be in awkward time because of the factors that I had described above (higher competition, software stops meaning just as much software as people would be custom-making them.)
I think that (former fly.io's) @tptacek's article[1] starts making more sense if viewed from this direction: What even is an OS now.
I don't have the answer to this question as to what happens next but its a form of development that I would prefer not to happen on a more gut instinct level?
Letting AI basically control everything and us not having any mental understanding of sorts and sort of becoming the meat-proxies just for economical reasons seems realistic possibility but a bleaker reality at that. I am left feeling a little bit uncomfortable if this reality turns out to be true.
[0]: https://news.ycombinator.com/item?id=49808422
[1]: https://sockpuppet.org/blog/2026/09/25/what-even-is-an-os-no...
The proof of the pudding.
OTOH, in the daily job, I have the team plan that's similar to 5x plan and I never had any limit problems, because I really need to understand be able to take responsibility for the code.
Totally different uses.
How you do this (and how deeply) I think is really the limit. I am doing this by focusing heavily on the design phase with grilling and trying to continually improve process to need less effort in the review phase. Are your models doing automated reviewing and testing before pushing out the PR (themselves)?
I think in the long run as models and the tools around them get better and cheaper, those that abdicate understanding will be able to achieve more. Although programmers think of that as irresponsible, ask yourself what does a tech lead do? And then what does a CTO do, etc?
2nd only to Opus 5.5, which is perfect. https://jonclegg.github.io/pacman-bakeoff/entries/claude-opu...
Up until very recently, all models struggled with this.
All results: https://jonclegg.github.io/pacman-bakeoff/
Could it be because the model was somehow pre-trained? If we compare it with pelicans that are still not-perfect…
Fable 5.1 was pretty good. Even animating it:
Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.
[1] Section 8.5 of the Sonnet 5.5 System Card
I was merely thinking of the theoretical aspect of it: performance of opus 5.5 is better than sonnet 5.5 across the board, with the exception of Terminal-Bench. So I was curious why this one stood out. Was it because they focused on it during training? Did sonnet 5.5 had access to more references for this benchmark? But based on my first reading, I concluded that it might just be the safety constraints that made the difference here, and I wanted to share that.
Anthropic made it that way, and I'd say the lower score is accurate.
> Claude Opus 5.5 scored 66.36% on Terminal-Bench 4.0 with safeguards enabled; requests flagged by the safeguards were answered by a fallback model following the default server-side fallback policy (2.5% of requests, affecting 10% of trials).
> https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba50242199...
I cannot find a Sonnet 5.5 system card.
AA intelegence index (agent harness doesn't have sonnet data yet) on max: Astra 27k Fable 5.1 78k (Sonnet 5) 118k Opus 5.5 119k Sonnet 5.5 193k
Opus 5 was previous record holder so hats off to Anthropic on blowing it away on token churn.
GLM and DeepSeek are great examples. They’re a bit like Linux or Android in that there isn’t necessarily one best provider. You need to do some research, try a few, and pick whatever works best for your use case.
I think that’s partly why Anthropic has been pushing its most expensive models so heavily for a while now. Sonnet and Haiku were great, but at that level of intelligence it’s becoming much harder for them to compete on price with Chinese models that have largely caught up.
The main reason to use frontier models from Anthropic or OpenAI now is the combination of intelligence and speed. Chinese frontier models still struggle to match that, possibly in part because of hardware constraints. But judging by the recent GLM releases, they seem to be moving in the right direction.
Most enterprise customers are paying per token at this point afaik, whether that’s to gh copilot, Anthropic, or running models on Vertex/Azure/Whatever
There's something to say for flat pricing rather than per token. Even if its not a better deal.
Read the full thread on Hacker News →
Related stories
- Hacker News · 4 points · 12 days ago
- Claude Sonnet 5.5 (Max Effort) Intelligence, Performance and Price Analysisartificialanalysis.aiHacker News · 4 points · 2 days ago
- Hacker News · 2 points · 7 days ago
- Hacker News · 1 points · 8 days ago
- Hacker News · 1 points · 11 days ago
- Claude Sonnet 5.5anthropic.comHacker News · 20 points · 2 days ago