Jev is a useful zero-shot classifier, but its probabilities can't be calibrated for your data. Calibration depends on your data distribution, which Jev never sees, so treat its outputs as scores and recalibrate them on…

65 points•alexmolas•8 days ago•61 comments•

61 comments

kantahayashi7 days ago
I tested Jev with a fair die 400 times without telling it the die result. The true probability of face 1 is 1/6, but Jev always chose face 1 and the probability it returned was about 83%. I also tested with a fair coin 200 times and got 0.92 probability.

I did several tests and I think Jev is good at problems with a correct answer but weak at problems about actual probabilities whose answers can't be known at all.

Write-up: "Jev Does Not Play Dice" https://kantahayashiai.github.io/posts/jev-does-not-play-dic...

alexmolas7 days ago
But "problems about actual probabilities whose answers can't be known at all" are exactly the problems where calibration is important. Since calibration is one of the big claims about Jev I'd expect it to perform well in these problems.
kantahayashi7 days ago
I agree. I think it's odd behavior too. Jev should be good at actual probability problems given the phrase "calibrated probabilities" TypeSafe uses for Jev. Maybe the reason is the data used in their training method (RLCD). If all the data consists of problems with a correct answer, I think this kind of odd behavior could happen.
throwaway_72747 days ago
If you instead offer probabilities as answers, it picks the right one with high credence.
drtz7 days ago
In the early Gemini 2 days (don't remember which version exactly) I had Gemini running as a voice assistant in my kitchen, and asked it to flip a coin and tell me if it was heads or tails. It responded with "heads". I was curious if it was actually doing something to simulate randomness, so I asked a few more times and saw a pattern: "tails", "heads", "tails", "heads"...

It continued alternating between the two until I got bored (around a dozen turns).

Unless your specific test is baked into its training, real probabilities require math and rough approximation at a minimum needs reasoning to sanity-check. Jev does neither. This isn't a new problem or anything unique to Jev.

tomrod7 days ago
The value of grandparent comment is that it identifies an edge case to keep in mind and make well-defined -- keeps us from blindly trusting.
DuperPower7 days ago
ask jev if god exists
edot7 days ago
Hah! I did the exact same tests as you! I found that if you give it the choice to say "not sure", it picks that 100% of the time. But if you pin it in a corner, then yes it does these weird things. Also yes, the continuous options were much more accurate than the choices. Not sure why that is.
tomrod7 days ago
Echoes a bit of a philosophical distinction with a long history: "Knightian Uncertainty" versus "Probability".
abhgh7 days ago
I like this post. I haven't had time to dig into Jev (they aren't accepting new signups), but calibrated probabilities is one of their pitches that caught my attention. And I was wondering how does one offer them on user data. Standard calibration essentially ensures that if a score of 0.8 accompanies a positive prediction (assuming the simple case of binary classification), then if you gathered together all predictions with a score of 0.8, around 80% will be correct.

If you have just one example you're sending to a model, how would they guarantee 80% over your data?

FYI, for an overview, scikit's page on calibration is great [1], and my answer on Quora from a long time ago covers a specific type [2].

[1] https://scikit-learn.org/stable/modules/calibration.html

[2] https://www.quora.com/How-is-isotonic-regression-used-in-pra...

danielmarkbruce7 days ago
While I don't believe they are doing the following: you can calibrate by inspecting the reasoning traces. That is the relevant distribution. If you ask someone to explain how/why they are classifying something one way v another, you can get a reasonably good understanding of their confidence level.
abhgh7 days ago
This tells me the confidence of the LLM's belief about the response - which is different from the calibrated confidence score. The former also is useful (just not what I thought their advertisement sells - and from the article it seems like it tripped up others as well), and there are different techniques to extract such a value [1] [2], typically via "response sampling", i.e., interrogate the LLM slightly differently to see if it changes its answer.

[1] Semantic Entropy https://www.nature.com/articles/s41586-024-07421-0

[2] Kernel Language Entropy https://openreview.net/pdf?id=j2wCrWmgMX

edot7 days ago
It's on OpenRouter if you want to try it.
abhgh7 days ago
Thank you!
jackb40407 days ago
This is why I don't understand why everyone's freaking out about it. By far the biggest problem with LLM classifiers is that they treat every individual business as the blurry average of all businesses in their training data. Being lighter is fine if you control for everything else, but at least at my company we would actually have room for a significantly more expensive / slower classifier if it were demonstrably better at following instructions.
bnbn887 days ago
This hype is caused by the price and the speed since most people don't know about small fast models and use big models for everything.
jrochkind17 days ago
They seem to have a really good social media astroturf marketing campaign.
softwaredoug7 days ago
It’s not just price and speed, the API is very well designed for classification. Going beyond the current structured outputs.
Tostino7 days ago
The available API shapes have limited so much over the years. It's really hard to come up with a new one and have it used, so everyone just tries to fit their work into the existing APIs (chat/completions).

Good job to them for putting out something that does seem quite nice to use, and will likely get a bit of wider traction.

croemer7 days ago
Article says that it _can_ be calibrated, it just isn't out of the box. So the title seems contradicted by the body.

> If you want calibrated probabilities you’ll still need to recalibrate Jev’s probabilities on your own data. The good news is that is cheap. A few hundred labeled examples from your actual data can be enough to fit a Platt scaling on top of Jev’s scores.

Groxx7 days ago
I suspect they meant it as "they claim it is already calibrated, but it can't be, because there is no universally correct calibration", and not "it is not possible to calibrate Jev", but I read it as the latter at first too.
alexmolas7 days ago
Yes, good point. I meant that the Jev can't be calibrated for everyone out of the box. I find it important because this is one of their main claims. They literally say

> Outcomes assigned a probability of 0.2 should occur about 20% of the time.

and this is not true

croemer4 days ago
Ah, I see, you mean "Jev cannot possibly be calibrated for everyone as there is always missing context".

I understood it as "It is impossible to calibrate Jev".

Read the full thread on Hacker News →

Related stories