Frontier models can now produce novel historical knowledge, but in a really weird way
2 comments
> They are also notably bad at judging the historical significance of what they find.
I use LLMs for some things that are outside the more common use-cases (in my case 3D design for 3D printing) and one thing I've noticed is that the errors it makes are so completely unlike human errors that they are hard to anticipate.
It will do things like build perfect snap catches but put them so the the pieces they are connecting are rotated 90 degrees from how they should be. It's "dumb" error, but hard to say the model itself if dumb because it does other very hard things so perfectly.
> seven chord groups
This sounds a lot more like Opus 5.0 than Opus 5.5 TBH. I wonder if that was an earlier investigation because 5.5 has improved that kind of language a lot.
You say that’s an error a human couldn’t do, but imagine if the human has never seen or touched the kind of item you were making and relied entirely on text descriptions to build its ontology. Off by 90 seems like such a believable mistake.
Maybe the model isn't intelligence in any form, except perhaps as an imperfect reflection of the intelligence of its training data.
Read the full thread on Hacker News →
Related stories
- Hacker News · 14 points · about 7 hours ago
- Hacker News · 1 points · 2 days ago
- The Verge · 0 points · 9 days ago
- Hacker News · 3 points · about 8 hours ago
- Claude Opus 5.5anthropic.comHacker News · 1788 points · 9 days ago
- Claude Opus 5.5 uses 95% fewer em dashes, but its answers are getting longerbleepingcomputer.comHacker News · 1 points · 4 days ago