Benchmark of TypeSafe's Jev against Sonnet 5, GPT-5 nano and local LLMs on 770 Reddit AITA verdicts: Brier scores, latency and cost - dchristopoulos/jev-aita

23 points•dchristopoulos•7 days ago•5 comments•

5 comments

edmundsauto7 days ago
It’s definitely not in the community’s interest to have auto posting bots also commenting, and I am not advocating for that

But in this case I appreciated the summary so i didn’t have to waste any time on the article.

Hopefully next time, my agent will read it all for me and make this comment.

maplet7 days ago
Using Reddit as the ground truth here feels like evaluating on training data
chaoz_7 days ago
Too much text. The only question I have and want to see at the top is whether it's more human aligned than LLMs (or, potentially, overfit)
dchristopoulos7 days ago
TypeSafe's new model, Jev, doesn't write text. You give it a situation and a question, and one quick call returns a probability for each answer.

The test: 770 posts from Reddit's r/AmItheAsshole. Each model had to predict the verdict Reddit actually gave.

Jev came second of seven setups. Sonnet 5 was a bit more accurate, though the lead is borderline once you account for how many comparisons were made.

Jev's median call was 6.3× faster than Sonnet's. Fast, but not 40–200×. That was measured on one laptop in one evening, so treat the exact number with some caution.

I have no affiliation with TypeSafe. Code, logs and method are open, so you can rerun it yourself.

dang6 days ago
Can you please not post AI-generated or AI-edited text to HN? It's not allowed here - see https://news.ycombinator.com/newsguidelines.html#generated and https://news.ycombinator.com/item?id=47340079.

Of course, it's impossible to know for sure what was LLM processed or not, but some of your posts (like this one) have been getting classified that way, and it looks like readers took it this way as well (https://news.ycombinator.com/item?id=49825833).

Read the full thread on Hacker News →

Related stories