Claude Opus 5.5 leads in agentic coding and knowledge work, and costs 40% less to run than Opus 5 on typical workloads.
1133 comments
Interesting how the very first line is used to remind the reader of their call to pace the frontier just last week, and everything else after that line is to demonstrate with very specific numbers how they absolutely are not pacing.
Edit: In response to the initial replies. To me it clearly means "releasing frontier models at any pace less than as fast as possible". It implies relative restraint compared to the previous state and without stating the degree of restraint.
https://www.war.gov/News/News-Stories/Article/Article/264106...
China is literally only a single step behind and willing to drop free models just to undercut the US companies.
I’m for it because I don’t want another massive Google or Meta.
I find it bizarre how intensely a bunch of these child/grandchild comments are criticizing the notion that people would even think to analyze the meaning behind the words.
Hacker News has always had a unique culture in which thoughtful discussion is basically the main goal, and it's intentionally incentivized in numerous ways. It's been my experience that any thoughts added to a post's conversation are seen as valuable as long as they are thoughtful and seeking to understand.
So these comments are clearly coming from a place that's antithetical to HN's culture. What that in mind, it seems likely to me (Occam's Razor) that these comments are either:
1. Astroturfing: Claude employees acting like everyday folks, secretly trying to shift public opinion.
2. AI cult mindset: "AI is humanity's salvation; how dare you have perspectives outside of those accepted by the cult."
Am I missing another likely option?
To bolster my point, right now we're posting on the top top-level comment, meaning a majority of active HN users find it to be a great addition to the conversation. Commenting to shut down the discussion is a red flag.
Prices per 1M tokens Claude Opus 5.5 Claude Opus 5
Cache reads $0.20 $0.50
Input tokens $4 $5
Output tokens $20 $25
Cache writes $5 $6.25
Opus 5 is the model with highest spend on openrouter (https://openrouter.ai/rankings#task-spend) and it seems plausible that Opus 5 is/was the highest spend model in the world, and certainly Anthropic's biggest moneymaker.If you are forced to reduce price despite raising capabilities, that certainly tells something about the market, and potentially about Anthropic future profitability too, since this model is their biggest topline contributor
Opus 5.5: Found 8/14 issues. Total cost: $15.40
Fable 5.1: Found 7/14 issues. Total cost: $66.34
Opus 5: Found 6/14 issues. Total cost: $15.19
Sonnet 5: Found 2/14 issues. Total cost: $19.15
This is a relatively small sample size, but it was both the best and the cheapest.
ETA: NB this is "Equivalent API" cost as reported by claude's CLI; I was using my subscription.
Fable 5.1 literally was a money grabber. While I liked the results, tokens were burned so hard it was embarrassing, while Astra seemed to not care.
Also Claude makes it very hard to pay for additional token budgets, allowing only credit cards. I don’t use mine anymore since I don’t need it in everyday life I was dumbfounded.
So Anthropic is just copying OpenAI so to say, matching them and essentially with Opus 5.5 being Fable 5.1 in disguise, all they do is reduce costs.
Competition works.
If 5.5 is any better, I might try to do agentic-assisted development instead of just telling fable to delegate
It doesn't look like that's happening, on the contrary the prices are falling especially when taking into account capabilities.
I'm hardly a fan of China/Xi, but I do appreciate and benefit from this.
In the end, everyone lost and there are millions of bikes in landfills.
If you're interested in the bikeshare bubble, Asianometry did a video on it a while ago.
I think this is what I'm most interested in. I mostly moved to Astra because I just can't work all day with the Claude Opus 5/Fable writing style. I don't think Astra is a better model, but it's the first OpenAI one that seemed good enough to me. Definitely keen to try Opus 5.5 and see if this claim is real.
Code-wise it seems to still nitpick, especially in reviews, but it doesn’t seem to rabbit hole quite as badly on tangents and scope-creep. These are just first impressions though. It’ll take a few weeks of regular use to really have a sense of it.
I hope Opus 5.5 is better, if for no other reason than all the Claude slop I have to read will be at least more tolerable.
One funny side effect of all of this: realizing that coworkers that use AI for almost all the text they generate at work have their writing style change every time a new model ships.
But oddly enough its still great at coding. Just like a lot of people it either interfaces well with people or machines but not both.
I'm good with DeepSeek v4.1 set to high. It is a relentlessly "hardworking" dirt cheap model.
Told it to convert a products page (that had two different fonts based on language) from two columns layout to 5 columns on desktop and 2 columns on mobile ensuring typography is readable.
My man went into spawning sub agent which failed to drive chrome so it wrote its own chrome driver protocol server in Typescript then generated a prototype website then downloaded the images and rendered each variation in a directory taking 100+ screenshots analyzing the typography depth and then delivering detailed report and then writing the whole thing with new page layout testing it again with several dozen screenshots using its driver and then saying all good and all really was good and whole thing took 25 minutes or so (including double visual validation) because it generates token at an incredible speed.
Total cost of the above? $0.07 cents.
PS: It generates token at such a blazing fast speed that you can't recognize the words as they are being added and can't read it without scrolling and pausing even if you're Jimmy Carter.
I’d not put Luna and Deepseek in the same tier as Sonnet, they were clearly ahead last time I checked (though I might be outdated and that’s on my personal use case).
And that all is 0.07 cents all included.
I was thinking of going with a subscription of Claude or Codex. The reason (at least that's what I am assuming): with OpenRouter or any PAYG per token setup there will be the anxiety of using up all the tokens in days or maybe 1-2 weeks instead of a month (say I set myself a budget of 15-20 USD per month, average equivalent of a usual subscription price).
Now I don't really want the top-notch models for the coding work I do.
So how much worth of "work/tokens" will I reasonably get for ≈$20 USD if I use it a lot? How much does that equal to - or is equivalent to, say in the world of subscription based Claude, Codex, or even GLM (with their 5-hour and all those cooldowns/limits)?
I am looking for a mental model/framework to visualise this. Can you (or anyone else reading this) please point me to a source where I can get some idea about this? I know I can just add $5 on OpenRouter and try to test. But I don't really know what/how to test these spends. I also want to understand how all this works. (I am new to agentic/llm world/coding, 2-3 months, after a career break of ~3 years, that too after working for more than a decade. I know, not at all good timing!)
Go to the "Cost" -> "Intelligence Index vs. Cost per Intelligence Index Task" That diagram maps their "Intelligence" score to "cost per task" and I think this gives a good basis on deciding with which model you want to go. Then you can either get an API token from that models provider directly or use openrouter and set openrouter to the model/providers of your choice.
You can also see on openrouter itself the details for each model like prices and what providers are offering it at what price.
Finally you can compare models details using openrouters compare feature like this:
https://openrouter.ai/compare/anthropic/claude-opus-5.5/open...
All four levels have a correctly shaped bicycle frame. The differences between the pelicans aren't huge, but the xhigh one has a better beak.
I haven't managed to get one for level "max" yet, it hit the limit of 128,000 cap for output tokens while it was still reasoning about the question!
Max started its thinking trace like this:
> This is a classic test request, so I want to plan out a well-composed pelican with its distinctive beak and pouch riding a bicycle with proper wheels, frame, and pedals, set against a simple sky and ground backdrop.
So that failed attempt on max cost me $2.56.
I ran this using my llm-anthropic plugin:
uv tool install llm
llm install llm-anthropic --upgrade
llm keys set anthropic
# paste key here
llm -m claude-opus-5.5 -o thinking_effort low "Generate an SVG of a pelican riding a bicycle"
# Then to save the markdown logs
llm logs -cu > logs-with-usage.mdIsn't this basically the model admitting it was trained on this? Otherwise why would it think a pelican svg is a usual request?
But its safe to say that pelicans on bicycles are disproportionally huge part of their training data
Doesn't mean Anthropic deliberately tried to train it to do a good job. If they DID train for the test their results are quite disappointing, I've seen better efforts from open weight Chinese models.
Off to a _great_ start...
Also interesting this somewhat mirrors my recent experience with Opus 5--too much effort and it starts looking for things to do and invents requirements that never existed
If you look carefully, everything except the last pelican has the two legs both in front of the crossbar as if the legs are all on one side of the bike.
The last pelican gets this correct.
Misplaced legs clearly indicate lack is spatial reasoning - the llm can reason about verbal idea of a bicycle but not about the actual object. The fact that this model got it correct gives me a pause. Did they figure out spatial reasoning? Or did this complain trickle down to the training set?
Fable 5.1 27 input, 65,927 output
Opus 5.5 27 input, 128,000 output (128k thinking tokens) - incomplete
[1] https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Read the full thread on Hacker News →
Related stories
- Claude Opus 5.5anthropic.comHacker News · 273 points · 8 days ago
- Hacker News · 2 points · 7 days ago
- DEV Community · 0 points · about 9 hours ago
- The Verge · 0 points · 8 days ago
- Hacker News · 4 points · 13 days ago
- Claude Opus 5.5 uses 95% fewer em dashes, but its answers are getting longerbleepingcomputer.comHacker News · 1 points · 3 days ago