Analysis of Anthropic's Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) and comparison to other AI models across key metrics including quality, price, performance (tokens per second & time to first…
106 comments
I've failed twice to get "Generate an SVG of a pelican riding a bicycle" to work with max, because in both cases it ran out of the 128,000 token budget while it was still reasoning about the problem.
I'm suspicious that "max" may be virtually useless if it's that easy to have it overthink to the point that it doesn't get to a response.
Transcript for one attempt here - expand the "Reasoning trace" bit to see it: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
I also switch to a better model for more complex tasks, also in low settings
I know there's been discussion about whether pelicanmaxxing is happening, but this is at least evidence that Claude was explicitly exposed to this problem.
My most exciting recent release is actually 5.6 Luna, not because it is the best on any index, but the dollar per work is insane value for money. I find myself more exciting by "value" than hypothetical ceilings because I'm just not in that budget category.
GPT-5.6 Sol's performance in the API should not change over time. If it has, that's a severe bug and we'll look into it.
We do sometimes tweak ChatGPT settings (e.g., tools, system prompts, efforts) over time, but we never play games to juice evals at launch times. You should always get what's advertised.
(I work at OpenAI.)
Yesterday, I ran an identical bug identification dataset from two weeks ago, saw a 50% drop from a few weeks ago, putting Sol on the same level as Luna. Sol had been finding 40-50 bugs per set, then dropped to 25, matching Luna’s performance. Not enough to establish a pattern, but enough to raise eyebrows.
Our review workflow is public if you want to peruse it, dataset isn’t. The process isn’t really stabilized yet either as I have to balance running this against limited budgets.
https://github.com/BiggerPockets/.github/blob/main/.github/w...
Moreover, the tests should be randomized somehow to ensure the models don't memorize the answer.
Edit: https://artificialanalysis.ai/models/claude-opus-5-5?models=...
It's also less clear what a lot of their metrics mean. Does Cost per Task include only things that can be verified to work and passed? As best I can tell, it does not.
I'm less concerned if one model's cost per task is $0.10 and another model's cost is $1.50 if the $0.10 task got it right 1% of the time and the $1.50 model got it right 66% of the time.
An equalized / weighted cost/time per task is much more valuable - being massively penalized for taking a lot of time and ultimately not passing when OTHER models did pass.
Not sure about how adaptive reasoning works though as they mention adaptive reasoning for every reasoning level
There are no prizes to be won by having the best model that’s 100x the price of something that’s good enough for 99% focuses cases.
Read the full thread on Hacker News →
Related stories
- Claude Opus 5.5 (High Effort) Intelligence, Performance and Price Analysisartificialanalysis.aiHacker News · 1 points · 8 days ago
- Claude Sonnet 5.5 (Max Effort) Intelligence, Performance and Price Analysisartificialanalysis.aiHacker News · 4 points · 2 days ago
- GPT-6.1 Sol (Max): Intelligence, Performance and Price Analysisartificialanalysis.aiHacker News · 4 points · 1 day ago
- GPT-6 Sol (Max) Intelligence, Performance and Price Analysisartificialanalysis.aiHacker News · 2 points · 8 days ago
- GPT-6 Luna (Max) Intelligence, Performance and Price Analysisartificialanalysis.aiHacker News · 2 points · 8 days ago
- Grok 4.7 Intelligence, Performance and Price Analysisartificialanalysis.aiHacker News · 5 points · 9 days ago