Benchmark of TypeSafe's Jev against Sonnet 5, GPT-5 nano and local LLMs on 770 Reddit AITA verdicts: Brier scores, latency and cost - dchristopoulos/jev-aita
5 comments
But in this case I appreciated the summary so i didn’t have to waste any time on the article.
Hopefully next time, my agent will read it all for me and make this comment.
The test: 770 posts from Reddit's r/AmItheAsshole. Each model had to predict the verdict Reddit actually gave.
Jev came second of seven setups. Sonnet 5 was a bit more accurate, though the lead is borderline once you account for how many comparisons were made.
Jev's median call was 6.3× faster than Sonnet's. Fast, but not 40–200×. That was measured on one laptop in one evening, so treat the exact number with some caution.
I have no affiliation with TypeSafe. Code, logs and method are open, so you can rerun it yourself.
Of course, it's impossible to know for sure what was LLM processed or not, but some of your posts (like this one) have been getting classified that way, and it looks like readers took it this way as well (https://news.ycombinator.com/item?id=49825833).
Read the full thread on Hacker News →
Related stories
- OpenAI says planned GPT-6.1 is too insecure to releasearstechnica.comArs Technica · 0 points · 1 day ago
- Hacker News · 1 points · 9 days ago
- Hacker News · 4 points · 6 days ago
- Hacker News · 2 points · 7 days ago
- Hacker News · 1 points · 5 days ago
- Hacker News · 3 points · 3 days ago