Claude Opus 5.5 leads in agentic coding and knowledge work, and costs 40% less to run than Opus 5 on typical workloads.

1806 points•km144•8 days ago•1133 comments•

1133 comments

sailingparrot8 days ago
> Claude Opus 5.5 is our first release since we called for pacing the frontier.

Interesting how the very first line is used to remind the reader of their call to pace the frontier just last week, and everything else after that line is to demonstrate with very specific numbers how they absolutely are not pacing.

mukmuk8 days ago
“Pacing the frontier” sounds smarmy and weird, like the phrase was generated by Claude itself
DiggyJohnson8 days ago
I really don't think it's productive for internet forums to constantly be criticizing language choice when the meaning is clear. Better to respond to the substance of the issue than word choice.

Edit: In response to the initial replies. To me it clearly means "releasing frontier models at any pace less than as fast as possible". It implies relative restraint compared to the previous state and without stating the degree of restraint.

tclancy8 days ago
It combines the elegance of LinkedIn-speak with the humbleness of desk-bound people who speak in military metaphor.
nradov8 days ago
It's interesting to compare Anthropic's language with what's coming out of the US military lately. They explicitly refer to China as the "pacing threat", meaning that there is a risk of China achieving superior military capabilities and thus we need to press forward with an arms race (including militarized LLMs) as fast as possible.

https://www.war.gov/News/News-Stories/Article/Article/264106...

topbanana8 days ago
You're right to call that out
nonethewiser8 days ago
IDK I think Anthropic is plenty smarmy and weird itself. Sounds like they wrote it.
apitman8 days ago
Isn't the idea that they claim to be willing to slow down if everyone does (ie governments force everyone to), but otherwise they won't slow down because they still think they'll make the best choices with superintelligence if they get there first? That's my understanding of what all the major labs claim to believe anyway.
johnfn8 days ago
Did no one in this entire HN thread read past the headline of the post from Amodei? He wrote a very clear set of actions Anthropic is taking.
SV_BubbleTime8 days ago
It’s just another cry for regulatory capture which imo is the only way they stay afloat at their current direction.

China is literally only a single step behind and willing to drop free models just to undercut the US companies.

I’m for it because I don’t want another massive Google or Meta.

dmazin8 days ago
It seems like they are. I mean, this is similar in performance to Fable (ish). It seems like more focus on making existing capabilities more accessible.
sailingparrot8 days ago
Fable 5.1 came out just 21 days ago. Only 3 weeks! And this is 20% relative improvement on terminal bench vs Fable 5.1 at less than half the price, and more human sounding output. does not feel paced to me tbh.
the_gipsy8 days ago
Occam's razor: they couldn't make any more substantial improvements.
lwhi8 days ago
And here we have it: the real reason that regulation has been called for, the ability to put out models that don't vastly outperform those previous, while not upsetting (future) shareholders.
bko8 days ago
I think the real reason is basic collusion. They're burning tens of billions training new models. It's a very competitive space. They know that the companies capable of frontier models is limited so if they can all get together and agree to slow down it gives them more time to make money from inference.
tencentshill8 days ago
What an amazing excuse for lower than expected performance! Our models are slow because we're so ethical.
PaulStatezny8 days ago
Thanks for spelling out what the original comment was implying.

I find it bizarre how intensely a bunch of these child/grandchild comments are criticizing the notion that people would even think to analyze the meaning behind the words.

Hacker News has always had a unique culture in which thoughtful discussion is basically the main goal, and it's intentionally incentivized in numerous ways. It's been my experience that any thoughts added to a post's conversation are seen as valuable as long as they are thoughtful and seeking to understand.

So these comments are clearly coming from a place that's antithetical to HN's culture. What that in mind, it seems likely to me (Occam's Razor) that these comments are either:

1. Astroturfing: Claude employees acting like everyday folks, secretly trying to shift public opinion.

2. AI cult mindset: "AI is humanity's salvation; how dare you have perspectives outside of those accepted by the cult."

Am I missing another likely option?

To bolster my point, right now we're posting on the top top-level comment, meaning a majority of active HN users find it to be a great addition to the conversation. Commenting to shut down the discussion is a red flag.

felixgallo8 days ago
in what way is beating every other frontier model with their own second-tier model, 'lower than expected performance'? Please be specific.
GodelNumbering8 days ago
Finally that price drop

   Prices per 1M tokens     Claude Opus 5.5    Claude Opus 5
   Cache reads              $0.20              $0.50
   Input tokens             $4                 $5
   Output tokens            $20                $25
   Cache writes             $5                 $6.25

Opus 5 is the model with highest spend on openrouter (https://openrouter.ai/rankings#task-spend) and it seems plausible that Opus 5 is/was the highest spend model in the world, and certainly Anthropic's biggest moneymaker.

If you are forced to reduce price despite raising capabilities, that certainly tells something about the market, and potentially about Anthropic future profitability too, since this model is their biggest topline contributor

kphorn8 days ago
I disagree - Fable melts the GPUs and they have a high incentive to move people off of that. If they have meaningfully decreased cost to serve on Opus 5.5, they can reduce prices and increase margin or at least turn off the most expensive compute.
AJ0078 days ago
It is only a price drop if price * tokens used is less
gwd8 days ago
Happened to be testing a "review patches on a mailing list" harness I was developing; here are a sample of the latest results, testing 12 patches containing a total of 14 issues:

Opus 5.5: Found 8/14 issues. Total cost: $15.40

Fable 5.1: Found 7/14 issues. Total cost: $66.34

Opus 5: Found 6/14 issues. Total cost: $15.19

Sonnet 5: Found 2/14 issues. Total cost: $19.15

This is a relatively small sample size, but it was both the best and the cheapest.

ETA: NB this is "Equivalent API" cost as reported by claude's CLI; I was using my subscription.

mcintyre19948 days ago
They're claiming a drop in token use too, and that it nets to 40% cheaper.
_the_inflator8 days ago
Claude adapts to OpenAI’s surprising move to simply deliver better performance than Fable 5.1, better tools as well as featuring very low pricing.

Fable 5.1 literally was a money grabber. While I liked the results, tokens were burned so hard it was embarrassing, while Astra seemed to not care.

Also Claude makes it very hard to pay for additional token budgets, allowing only credit cards. I don’t use mine anymore since I don’t need it in everyday life I was dumbfounded.

So Anthropic is just copying OpenAI so to say, matching them and essentially with Opus 5.5 being Fable 5.1 in disguise, all they do is reduce costs.

Competition works.

blfr8 days ago
People are paying for Opus 5? Not just burning down tokens left after they enjoyed Fable on the sub? Amazing.
rapfaria8 days ago
My workplace doesn't even offer Fable. And on the sub, I've had a hard time understanding Opus 5, but Fable can deal with it with subagents.

If 5.5 is any better, I might try to do agentic-assisted development instead of just telling fable to delegate

herpdyderp8 days ago
When you need to disable data retention, you cannot use subscription plans.
zanderwohl8 days ago
Fable tends not to perform better, just cost more.
neuronexmachina8 days ago
Enterprise and most Team accounts use API pricing, they don't have an included-usage quota.
girvo8 days ago
We can’t use fable at work, opus and Astra are as good as it gets.
margorczynski8 days ago
All the anti-AI people constantly say that any moment now the prices will skyrocket and in the end human work will be cheaper compared to using AI.

It doesn't look like that's happening, on the contrary the prices are falling especially when taking into account capabilities.

johnecheck8 days ago
It's the Chinese open source models. They're barely behind the frontier, making AI a commodity, forcing openAI and Anthropic's margins downward.

I'm hardly a fan of China/Xi, but I do appreciate and benefit from this.

hajile8 days ago
These companies are posting massive losses while also lowering prices. This sounds just like the Chinese bikeshare bubble where they were all taking massive losses in hopes that their competitor would go broke first.

In the end, everyone lost and there are millions of bikes in landfills.

If you're interested in the bikeshare bubble, Asianometry did a video on it a while ago.

https://www.youtube.com/watch?v=FQrEDq8KPiU

runtime_terror7 days ago
You have no suspicion that this most recent collusion is in part to setup the conditions to guarantee government subsidies/bailouts?
dboreham8 days ago
There's zero chance of that ever happening. Pure delusion.
mcintyre19948 days ago
> Communication. Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5. It puts the most important information up front, and its style makes it a better work partner over long sessions. As one early tester put it, “it writes the way I do.” In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one.

I think this is what I'm most interested in. I mostly moved to Astra because I just can't work all day with the Claude Opus 5/Fable writing style. I don't think Astra is a better model, but it's the first OpenAI one that seemed good enough to me. Definitely keen to try Opus 5.5 and see if this claim is real.

ascendantlogic8 days ago
So far it seems the same. I used Opus 5.5 for an hour this evening and it was just as painfully verbose as Opus 5. It also used the term "load bearing" 4 separate times.
bobbylarrybobby8 days ago
I noticed that when opus 5.5 was on parts on my codebase that had lots of 5.0-generated comments, it picked up its style. Unfortunately I think 5.5 has been trained to mimic what it sees so that you can ease up on the instructions, but this does mean parts of your codebase that 5 touched will be somewhat viral.
loveparade8 days ago
I will forgive it using load bearing as long as it finds the right seams.
therealdrag08 days ago
Due to this release note, I used it once to rewrite a doc, and was disappointed.
MattyRad8 days ago
If anything it seems worse. I'm experiencing about -10% insufferable jargon, but +30% more verbosity. It's unredeemable. There also seems to be even less structured output (headings, bullets, etc).
samat7 days ago
thank you, i did not try it myself and having read this, will not waste my time and energy
invalidusernam37 days ago
I would happily pay more for a claude model that performs the same but speaks normal English. The proliferation of claudespeak in the workplace is driving me insane. Every ticket, every PR feedback, every comment in the codebase is poisoned with its ridiculous unnatural vocabulary
sdthjbvuiiijbb8 days ago
I'm surprised that you're getting so many replies saying it's the same. So far in my usage today Opus 5.5 does seem like a noticeably better writer. Opus 5 frequently made me want to strangle it while 5.5 has been producing a lot less incomprehensible gobbledygook.
MattyRad8 days ago
I know we're all experiencing NDFSMs differently, like that's part of the whole problem, but 5.5 just gave me "The truncating quantizer collides two oranges", which is a new low for me.
pixelready8 days ago
Yeah I’m having a much better time reading Opus 5.5 output today vs. 5’s wall of nonsense. You still get a few telltale turn of phrases, though the load-bearing smoking guns haven’t turned up yet. It’s still a bit verbose compared to what I’d ideally like, but it’s tolerable now.

Code-wise it seems to still nitpick, especially in reviews, but it doesn’t seem to rabbit hole quite as badly on tangents and scope-creep. These are just first impressions though. It’ll take a few weeks of regular use to really have a sense of it.

derangedHorse8 days ago
As someone who uses both, Astra was 100% the better model. I have yet to give 5.5 a spin so maybe that’ll be the new top contender.
BatFastard8 days ago
I prefer Astra for creative uses, Fable seems better for hardcore coding.
Trasmatta8 days ago
Opus 5 has made me question my sanity on a daily basis, especially as all my coworkers started lobbing Opus 5 slop grenades everywhere. It had the worst and most infuriating writing style I've ever seen.

I hope Opus 5.5 is better, if for no other reason than all the Claude slop I have to read will be at least more tolerable.

One funny side effect of all of this: realizing that coworkers that use AI for almost all the text they generate at work have their writing style change every time a new model ships.

nonethewiser8 days ago
I really wonder how it converged on its style. It's pretty unique and terrible. It's not like it's just mimicking something or it was purposefully design to be that way. I mean the reason may be diffuse and uninteresting... just the result of a lot of factors and lack of control over the writing style probably.

But oddly enough its still great at coding. Just like a lot of people it either interfaces well with people or machines but not both.

atombender8 days ago
Astra is better here, but the one I'm the most impressed with is Gemini. It's always been good, but 3.6 Flash is even better. It writes in a pleasant, human style. Not perfect, but it has a good balance between technical accuracy and readability that is better than what I've seen from any other mainstream model.
LtdJorge8 days ago
Yes, it made me want to vomit. If the new Fable only changed the writing style to just sound like a human, same performance for everything else, I'd be pretty happy.
rfgplk8 days ago
Opus is only usable if you have a post-turn formatter that strips all comments from the generated source. I'm not even kidding it's that bad.
Aperocky8 days ago
It's not X, it's Y, not A, not B, not C, and he haven't even woken up yet! Here's the catch, the detail is in the devils and the twist is that it's designed!
wg08 days ago
No thanks.

I'm good with DeepSeek v4.1 set to high. It is a relentlessly "hardworking" dirt cheap model.

Told it to convert a products page (that had two different fonts based on language) from two columns layout to 5 columns on desktop and 2 columns on mobile ensuring typography is readable.

My man went into spawning sub agent which failed to drive chrome so it wrote its own chrome driver protocol server in Typescript then generated a prototype website then downloaded the images and rendered each variation in a directory taking 100+ screenshots analyzing the typography depth and then delivering detailed report and then writing the whole thing with new page layout testing it again with several dozen screenshots using its driver and then saying all good and all really was good and whole thing took 25 minutes or so (including double visual validation) because it generates token at an incredible speed.

Total cost of the above? $0.07 cents.

PS: It generates token at such a blazing fast speed that you can't recognize the words as they are being added and can't read it without scrolling and pausing even if you're Jimmy Carter.

some_infra_dude8 days ago
What are you working on? Im always a bit surprised by folks that seem ok with non frontier models - the quality is just not there. I've found most code produced by even luna / sonnet tier models to be significantly worse quality. It seems to me they can't handle any mild complexity at all. Are you just prompting very explicitly and detailed?
knapcio8 days ago
I’ve been working with frontier models for the last year on a large real-world project with multiple apps. Eventually they started becoming unreliable. Simple UI issues, usage limits and overengineering became constant problems. I had to come up with a robust process: task analysis, user intent analysis, implementation and testing. All of these steps loop when needed with 15-20 steps per task in total. This finally got the models to do a good job but I started hitting usage limits. Then I switched to DeepSeek and haven’t noticed any drop in quality. It’s fast at around 300 TPS and cheap. I don’t miss frontier models anymore! Come up with a good process and you might not need them either.
jwrallie8 days ago
Yes, I tell exactly how I want it done and what files to modify, in those conditions sometimes a model that does not try to read between the lines works better.

I’d not put Luna and Deepseek in the same tier as Sonnet, they were clearly ahead last time I checked (though I might be outdated and that’s on my personal use case).

anewhnaccount28 days ago
Deepseek on high reminds me of Opus 4.x which was quite good for a lot of tasks.
cbg08 days ago
If you're producing slop, the quality of the model is irrelevant as long as it compiles. I frequently catch Opus/Sol making silly errors and over-engineering solutions while small ones like Luna struggle with complex tasks.
glub8 days ago
Also include that all of this comes with full reasoning traces, so if something goes wrong, you know exactly what assumption it started from.
wg08 days ago
Yes exactly. Reading this "thinking" traces is a great tool.
soundworlds8 days ago
Same here! DeepSeek v4.1 Flash has been my moment of "does everything I need, cheaply. Please now focus all R+D on making this efficient enough to run off a laptop"
tontinton8 days ago
Competition is good
wg08 days ago
It really is good. I forgot to mention that within that said sub agent, it also went into exploring top e-commerce websites (Zalaondo, Temu, Amazon, eBay) for exploring prevailing industry UX best practices and taking screenshots of their product and category pages with its own written chrome driver that I talked about and then went onto prototyping a new website in a temporary directory and then taking hundreds of screenshots to analyse what would be the best column density one each medium for each language.

And that all is 0.07 cents all included.

shelled8 days ago
Hey. I am on a GLM Coding plan subscription (old price; their base coding plan) right now and share the key (this will go away soon).

I was thinking of going with a subscription of Claude or Codex. The reason (at least that's what I am assuming): with OpenRouter or any PAYG per token setup there will be the anxiety of using up all the tokens in days or maybe 1-2 weeks instead of a month (say I set myself a budget of 15-20 USD per month, average equivalent of a usual subscription price).

Now I don't really want the top-notch models for the coding work I do.

So how much worth of "work/tokens" will I reasonably get for ≈$20 USD if I use it a lot? How much does that equal to - or is equivalent to, say in the world of subscription based Claude, Codex, or even GLM (with their 5-hour and all those cooldowns/limits)?

I am looking for a mental model/framework to visualise this. Can you (or anyone else reading this) please point me to a source where I can get some idea about this? I know I can just add $5 on OpenRouter and try to test. But I don't really know what/how to test these spends. I also want to understand how all this works. (I am new to agentic/llm world/coding, 2-3 months, after a career break of ~3 years, that too after working for more than a decade. I know, not at all good timing!)

kristophph8 days ago
A good resource is this: https://artificialanalysis.ai/?cost=intelligence-vs-cost-per...

Go to the "Cost" -> "Intelligence Index vs. Cost per Intelligence Index Task" That diagram maps their "Intelligence" score to "cost per task" and I think this gives a good basis on deciding with which model you want to go. Then you can either get an API token from that models provider directly or use openrouter and set openrouter to the model/providers of your choice.

You can also see on openrouter itself the details for each model like prices and what providers are offering it at what price.

Finally you can compare models details using openrouters compare feature like this:

https://openrouter.ai/compare/anthropic/claude-opus-5.5/open...

simonw8 days ago
Here are pelicans for thinking levels low, medium, high, and xhigh: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

All four levels have a correctly shaped bicycle frame. The differences between the pelicans aren't huge, but the xhigh one has a better beak.

I haven't managed to get one for level "max" yet, it hit the limit of 128,000 cap for output tokens while it was still reasoning about the question!

Max started its thinking trace like this:

> This is a classic test request, so I want to plan out a well-composed pelican with its distinctive beak and pouch riding a bicycle with proper wheels, frame, and pedals, set against a simple sky and ground backdrop.

So that failed attempt on max cost me $2.56.

I ran this using my llm-anthropic plugin:

  uv tool install llm
  llm install llm-anthropic --upgrade
  llm keys set anthropic
  # paste key here

  llm -m claude-opus-5.5 -o thinking_effort low "Generate an SVG of a pelican riding a bicycle"

  # Then to save the markdown logs
  llm logs -cu > logs-with-usage.md
MikhailTal8 days ago
> This is a classic test request

Isn't this basically the model admitting it was trained on this? Otherwise why would it think a pelican svg is a usual request?

Brendinooo8 days ago
Plenty of times I’ve seen a model say “it’s a classic X” despite not being a classic anything. Might just recognize it’s a test in general, or it might just be a tic.
MaxikCZ8 days ago
Dont conflate "I know this is test case" with it being trained on it.

But its safe to say that pelicans on bicycles are disproportionally huge part of their training data

simonw8 days ago
It's the model admitting that it has heard of the test. It's been around for a couple of years now so I'd be surprised if it hadn't.

Doesn't mean Anthropic deliberately tried to train it to do a good job. If they DID train for the test their results are quite disappointing, I've seen better efforts from open weight Chinese models.

zamadatix8 days ago
I think people just like to see the drawings at this point.
FergusArgyll8 days ago
It has read the internet. That doesn't mean it was literally RL'ed for this
nijave8 days ago
>I haven't managed to get one for level "max" yet, it hit the limit of 128,000 cap for output tokens while it was still reasoning about the question!

Off to a _great_ start...

Also interesting this somewhat mirrors my recent experience with Opus 5--too much effort and it starts looking for things to do and invents requirements that never existed

ceroxylon8 days ago
I was a bit skeptical when they said it behaves like Fable but is cheaper... those two things have been mutually exclusive in my experience, no LLM can light tokens on fire faster while spinning its wheels than the Fable/Mythos tier of models.
adverbly8 days ago
> The differences between the pelicans aren't huge, but the xhigh one has a better beak.

If you look carefully, everything except the last pelican has the two legs both in front of the crossbar as if the legs are all on one side of the bike.

The last pelican gets this correct.

DenisM8 days ago
I’ve been paying attention at this exact detail.

Misplaced legs clearly indicate lack is spatial reasoning - the llm can reason about verbal idea of a bicycle but not about the actual object. The fact that this model got it correct gives me a pause. Did they figure out spatial reasoning? Or did this complain trickle down to the training set?

narmiouh8 days ago
It is interesting that Fable 5.1 max [1] which also produced a decent pelican with 65k output tokens compared to 5.5 running out of 128k output tokens tells us something about the new models token usage propensity despite this being a sample of 1.

Fable 5.1 27 input, 65,927 output

Opus 5.5 27 input, 128,000 output (128k thinking tokens) - incomplete

[1] https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

inshard8 days ago
Not as good as Astra or Fable 5.1 on this test as far as I can see. I wonder if any benchmark exists for artistic taste, visual sophistication etc. I think your Pelican test does touch on these aspects of a model and is useful for developers trying to build rich digital experiences (includes games, interactive websites and apps). These benchmarks are subjective so it may not be easily established and will have polarized reactions before it gains legitimacy. May even need human judgement layers adding to the cost of running it.
skerit8 days ago
I like the Pelican test. And I agree this pelican looks very boring.

Read the full thread on Hacker News →

Related stories