Sonnet 5.5 scores just behind Opus 5.5 on Artificial Analysis Intelligence Indexartificialanalysis.ai
Anthropic's new Sonnet model scores 56, just 2 points behind Opus 5.5 (max), but at the highest Output Tokens per Task we have measured
1 comment
simianwords2 days ago
It’s better than both fable and Astra? I’m going to start trusting this index less and less.
jug1 day ago
It's apparently great at Max but extremely costly. Looks best to me at Medium and High. Over that and I'd go Opus. I consider Fable largely obsolete.
simianwordsabout 21 hours ago
Astra at light is worse than sol at light tbh
ramon1562 days ago
the issue is, what will we use? companies are benchmaxxing, so how do you make a bench that cannot be cheated?
Human curated benches aren't accurate enough
epolanski1 day ago
The problem with benchmarks is that they are end to end.
You go from one prompt to the final solution, whereas for many of us it is about the experience of iterative, multi turn working.
Read the full thread on Hacker News →
Related stories
- The Verge · 0 points · 5 days ago
- Can John Ternus find Apple’s next big thing?theverge.comThe Verge · 0 points · 10 days ago
- The Verge · 0 points · 3 days ago
- The Verge · 0 points · 12 days ago
- I have some questions for Mark Zuckerbergtheverge.comThe Verge · 0 points · 7 days ago
- Sonnet 5.5 trumps Fable 5.1 on Artificial Analysisartificialanalysis.aiHacker News · 3 points · 2 days ago