Analysis of Anthropic's Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) and comparison to other AI models across key metrics including quality, price, performance (tokens per second & time to first…

333 points•theanonymousone•8 days ago•106 comments•

106 comments

simonw8 days ago
This is the page for the "max" reasoning setting. The page for xhigh is https://artificialanalysis.ai/models/claude-opus-5-5-xhigh and the page for medium (the default setting) is https://artificialanalysis.ai/models/claude-opus-5-5-medium

I've failed twice to get "Generate an SVG of a pelican riding a bicycle" to work with max, because in both cases it ran out of the 128,000 token budget while it was still reasoning about the problem.

I'm suspicious that "max" may be virtually useless if it's that easy to have it overthink to the point that it doesn't get to a response.

Transcript for one attempt here - expand the "Reasoning trace" bit to see it: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

zerof1l8 days ago
This is totally a thing I noticed myself about 3 months ago. Medium thinking effort is ideal for most tasks. At high and above, models tend to generate more output in the form of comments or code for the same problem with no real benefit. Its a self-feeding loop: more output becomes more input, which then becomes more output. High is the highest I go. If I need more intelligence, it's better to use a more powerful model with less thinking effort or break the problem into phases. Much better result.
zozbot2348 days ago
This version of Opus "max" apparently has even higher thinking output than Qwen "max", which is infamous for its thinking streams where it constantly second-guesses itself, then third-guesses, fourth-guesses and generally nth-guesses itself for arbitrarily large n. Of course, we aren't actually seeing Claude's raw thinking output: all we get is the after-the-fact prettified "summary". One wonders how much of that is a coincidence, or whether there's a reason behind that.
realusername8 days ago
Personally I use everything in low reasoning. Maybe I'm wrong but I think that the higher reasoning settings are almost never worth it, it's marginal gains for a much higher budget.

I also switch to a better model for more complex tasks, also in low settings

RGS18118 days ago
"This is a classic test request..."

I know there's been discussion about whether pelicanmaxxing is happening, but this is at least evidence that Claude was explicitly exposed to this problem.

croemer8 days ago
Of course it was exposed - not sure it's explicit or not. Why wouldn't HackerNews comments be part of the training data? And Simon's blog and the many discussions about Pelicans? It'd be hard to miss. Doesn't mean Anthropic has made this an explicit goal in training.
simonw8 days ago
See here for more discussion of that: https://news.ycombinator.com/item?id=49803892#49804881
pgwhalen8 days ago
It would be genuinely shocking at this point if any of the frontier models weren't well exposed to the problem.
cubefox8 days ago
The model recognizing the task doesn't mean it was benchmaxxed (RLVR-trained) to solve it. It might simply recognize it from pre-training on Internet text.
0x10ca1h0st8 days ago
Lets start frog riding motorcycle trend until they frogmaxx, or cat driving convertible.
Someone12348 days ago
For people with any kind of budget, Opus 5.5's [Medium] actually can make sense dollar per intelligence/dollar per task wise. Heck, it puts some other models to shame. [Max]'s cost is completely unhinged.

My most exciting recent release is actually 5.6 Luna, not because it is the best on any index, but the dollar per work is insane value for money. I find myself more exciting by "value" than hypothetical ceilings because I'm just not in that budget category.

seabass-salmon8 days ago
That was true for me four weeks ago, but 2-3 weeks ago Luna turned into drivel in essentially the same complexity of task. I feel it came back somewhat in recent days but does feel like it's being manipulated.
dc4438 days ago
i assume you mean 6 Luna, because we've had 5.6 luna for ages
Gcam8 days ago
Hey! From the Artificial Analysis team. We also have a model releases page which shows all reasoning efforts (not just max), including the trade off curves https://artificialanalysis.ai/models/releases/claude-opus-5-...
samuelknight8 days ago
I have experienced this with open weight models too. "Max" is for benchmaxxing the intelligence metric and is not meant for use in productive work. Like drawing pelicans.
breckenedge8 days ago
Do these evaluations get re run a few weeks after launch? I started doing that yesterday for our internal dataset and found Sol’s performance had regressed to be equal to Luna’s. Granted this was one run, but something I’m becoming more concerned about, the model providers want to quickly prove they’re the best, people switch to them, then they pull the rug.
tedsanders8 days ago
Can you share more information on your methodology?

GPT-5.6 Sol's performance in the API should not change over time. If it has, that's a severe bug and we'll look into it.

We do sometimes tweak ChatGPT settings (e.g., tools, system prompts, efforts) over time, but we never play games to juice evals at launch times. You should always get what's advertised.

(I work at OpenAI.)

breckenedge8 days ago
This is all code review runs via OpenRouter with a Pi harness, and it’s totally possible there are shenanigans going on elsewhere.

Yesterday, I ran an identical bug identification dataset from two weeks ago, saw a 50% drop from a few weeks ago, putting Sol on the same level as Luna. Sol had been finding 40-50 bugs per set, then dropped to 25, matching Luna’s performance. Not enough to establish a pattern, but enough to raise eyebrows.

Our review workflow is public if you want to peruse it, dataset isn’t. The process isn’t really stabilized yet either as I have to balance running this against limited budgets.

https://github.com/BiggerPockets/.github/blob/main/.github/w...

sunaurus8 days ago
Do you have any hypothesis for why this experience is so consistently reported by users (anecdotally)? Seemingly across all providers.
dr_kiszonka8 days ago
But the chat version does change over time, correct? It has been my experience that Sol's performance has deteriorated significantly.
doctorpangloss8 days ago
ArtificialAnalysis tweaks stuff until newest big proprietary model is on top, not you haha
mnicky8 days ago
Well there is at least the degradation tracker from Margin labs for Sol and Opus: https://marginlab.ai/trackers/codex/
therealdrag08 days ago
It looks pretty consistent? At least within reason for a stochastic model. Or am I missing something?
scrollop8 days ago
Need one of these for openai:

https://marginlab.ai/trackers/claude-code/

roblabla8 days ago
Wouldn't that be basically https://marginlab.ai/trackers/codex/ ?
echelon8 days ago
These tests need to be sampled continuously.

Moreover, the tests should be randomized somehow to ensure the models don't memorize the answer.

hglaser8 days ago
Half the cost per task compared to Opus 5, comparing high effort to high effort. That's just really nice.

Edit: https://artificialanalysis.ai/models/claude-opus-5-5?models=...

tomjakubowski8 days ago
Tasks are completed in about half the time too. Although we'll see if it slows down in a few weeks as Anthropic's model services are prone to do.
sharktheone8 days ago
That is a lot. I thought Anthropic models would just do the opposite because they are greedy for money.
giancarlostoro8 days ago
Greed is not what's driving these prices, its cost. They considered very much in the red.
user439288 days ago
Astra High is slightly cheaper at $1.73 vs $1.82 for Opus 5.5
onlyrealcuzzo8 days ago
The UI/UX seems impressively bad. DeepSWE's cost curve has a better, more obvious way to sort by only the top level of reasoning to avoid 80% of the graph just being the same 3-5 models at their 8 different reasoning levels...

It's also less clear what a lot of their metrics mean. Does Cost per Task include only things that can be verified to work and passed? As best I can tell, it does not.

I'm less concerned if one model's cost per task is $0.10 and another model's cost is $1.50 if the $0.10 task got it right 1% of the time and the $1.50 model got it right 66% of the time.

An equalized / weighted cost/time per task is much more valuable - being massively penalized for taking a lot of time and ultimately not passing when OTHER models did pass.

makeavish8 days ago
Nice catch, AA only shows max effort by default and I got disappointed thinking it's a token guzzler though: https://artificialanalysis.ai/models/claude-opus-5-5?models=...

Not sure about how adaptive reasoning works though as they mention adaptive reasoning for every reasoning level

cmiles88 days ago
This continues to show that these foundational models are only slightly better than open weight models but cost around 100x as much. The history of tech is riddled with “good enough” eating “best” for lunch all day long. Unless the big labs come up with a viable business plan pronto it’s looking like AI will be no different.

There are no prizes to be won by having the best model that’s 100x the price of something that’s good enough for 99% focuses cases.

conception8 days ago
Unless the 100% is super intelligence. Thats the bet being made with our economy/society.
antupis8 days ago
Pretty much this, I don't just see how ASI could happen with current LLMs without big theoretical breakthrough so that's why 'phasing the frontier' and 'distillation attacks'.
linuxrebe18 days ago
Fingers are crossed on this one. I had gone back to using opus 4.8 instead of using opus 5. Simply because 4.8 is much better at remembering what it's doing and following instructions than 5. 5 often had a tendency to get halfway through solving a problem and then I would have to stop it in the middle, because it had lost its way and was going off on a tangent rather than dealing with the problem. In that respect, 4.8 was a lot more stable.

Read the full thread on Hacker News →

Related stories