Everyone is talking about Jev - here it is in 25 lines of Python.

691 points•bashbjorn•8 days ago•212 comments•

212 comments

sigmoid108 days ago
Going directly for the logprobs is always icky when you use a chat model as base, because they are trained to write prose as output. So your "choice" tokens and thus their probabilities might get diluted in whatever else it wanted to say. If you have to do it in the same way as this post, at least add clear system instructions and a carefully worded beginning to the assistant output section of the prompt to lower the chances of it wandering off immediately.

I've found that using structured outputs solves this problem much better. Instead of letting a model generate only "A", "B" or "C" and looking at the probs, have it directly generate "Legitimate", "Spam" or "Phishing" or any other pre-defined option from a set of multi-token sequences. Behind the scenes it boils down to something quite similar, but you're not running into the risk that the model actually wanted to say "A phishing attempt seems likely, so answer (C) is correct.", which would lead "A" to have the highest probability in the first token. You can even use a reasoning budget this way either via inherent reasoning or a free-form part preceding the remaining output structure. You can also have it assign probabilities (either in words or numbers) using more complex output structures, but I would not rely on them much more than the token logprobs (they can still be quite good though).

dTal8 days ago
The whole point is the quantified output. If you just ask an LLM to type out its confidence "manually", it'll make up some nonsense. The logprob numbers are more reliable.

I got this technique to work extremely reliably last year. However there were a bunch of caveats: 1) Firstly, you must institute a check that the multiple choice tokens dominate the output distribution. They should sum to 95% or more, ideally 99%, or the LLM is not following instructions properly. This is also the problem with constrained decoding - if the LLM really doesn't want to output a valid answer, the one you extract will not be high quality. 2) You need to ask it multiple times, permuting which option corresponds to which letter, and average the results. LLMs are surprisingly biased towards picking "A", especially if they're otherwise not sure. 3) For the same reason, performance improves if you frame the prompt as if it were the middle of a quiz. "Question 1" carries baggage that "Question 12" doesn't. 4) You must be exceedingly careful with tokenization.

But when all was said and done, I got a general purpose A/B classifier that gave high resolution quantitative output for the cost of a couple dozen tokens ingested and a couple inference passes.

sigmoid108 days ago
>The whole point is the quantified output. If you just ask an LLM to type out its confidence "manually", it'll make up some nonsense. The logprob numbers are more reliable.

The whole point of my argument is that neither is good, but from a technical perspective logprobs is probably the worst unless you train a model on specific outputs. In which case you'd throw out the generality again, so when I think about it more, it's actually the worst overall. In my experiments, having the model simply assign "high" or "low" probability in a structured output generally performs best. You can try numbers, but you will never get anything close to what you could expect from traditional ML. And most certainly not from logprobs.

TeMPOraL8 days ago
> LLMs are surprisingly biased towards picking "A"

GP pointed at a causal explanation for this: almost every sentence in English that's a statement will start with "A" or "An", so "biased towards picking ''A''" will include most attempts at saying anything long-form for any reason.

boredumb8 days ago
> LLMs are surprisingly biased towards picking "A", especially if they're otherwise not sure.

Not nearly as sophisticated as myself who would mutter "When in doubt - Charlie out" before marking C.

dragonwriter7 days ago
Sure, it was high resolution (precise), how was accuracy compared to Jev (or existing open source implementations of the same concept, like laya)?

Also, Jev/laya do it in one forward pass, for multiple questions about the same state, rather than multiple passes for one question about that state. Well, for the usual multilingual configuration, two forward passes through different small models for laya, but that's because one is the router which chooses which model should do the real work, but still.

ainch8 days ago
In my experience as well using logprobs to try to quantify uncertainty, LLMs are a poor fit. Neural nets in general struggle with 'calibration' --- ie. if a prediction is truly 50/50, neural nets are often prone to predicting overconfidently [0].

I ran some tests using GPT-4 to do some basic classification a couple years ago. On ambiguous options which had to be escalated to a human, the LLM would regularly output something like a 99.8% probability, compared to 99.99% for a correct answer.

0: https://arxiv.org/pdf/1706.04599

nautilus508 days ago
+1, llama.cpp has a --grammar parameter which you can pass a BNF style grammar file to constrain generation. It can be used in Python llama.cpp wrapper

https://til.simonwillison.net/llms/llama-cpp-python-grammars

porridgeraisin8 days ago
Yes. But even then, the probabilities are not calibrated. In jev/laya, they are (well, relatively anyways).
foo12bar8 days ago
If we're talking about running it locally, what about passing a partial response as part of the input?

Prompt part: "What is better, toast or bread?"

Incomplete answer part: "The answer to this question is "

and then have the LLM finish the answer. I did this with subtitle translation using llama.cpp (with Python) and had great success. Just past 5 already translated subtitles as the incomplete answer, and the LLM infallibly just continues to translate. No markdown, and usually no talkback if the subtitles contain nasty subjects like bioweapons or nuclear stuff. It just works.

_davide_8 days ago
Agreed, it's a real issue, but it can probably be vastly reduced by having the schema in the system prompt and by giving the model an expectation of a fixed value: no decent modern would pick a prose ligament over a provided value.

To completely squash the issue, a few cheap LoRa iterations will do the trick just fine.

wongarsu8 days ago
Sure, you can fix that in a couple lines. Then a couple more lines for evaluating multiple questions on the same answer in parallel. Then a couple more lines for the confidence score (which is trivial to compute from all we have, but missing regardless). Then a harness to fine-tune an existing model to perform better on this specific task, and a collection of training data to use for that

I think we can all agree that Jev is not rocket science. It's a good idea executed well, with marketing that might have been a tad too bold

antirez8 days ago
Because of masked attention in LLMs, if you put the options before the body (the email to analyze), the transformer already knows what it needs to look for, and can use more tokens to create state to address that specific task (BERT has no mask in the attention, so tokens attend also to next tokens). You could also do a few examples in the system prompt to improve calibration.

Another trick that works is to repeat the question two times: "I'm repeating the task and labels for clarity: ..."

__jf__7 days ago
Wow! TIL! I've been running a for loop around the two ordering variations to catch the winner of each turn and the difference is quite noticeable. In the options-after-body case in 47 of 100 attempts it classifies as phishing, whereas in the options-before-body case it classifies clearly as rickroll (94 out of 100 attempts)

Payroll sends you an email with a link to a Youtube video that plays a song.

Options after body:

    Average probabilities:
    Rickroll   0.5158 ( 51 wins)
    Phishing   0.4561 ( 47 wins)
    Spam       0.0281 (  2 wins)
    Joke       0.0000 (  0 wins)
    Legitimate 0.0000 (  0 wins)

Options before body:

    Average probabilities:
    Rickroll   0.9293 ( 94 wins)
    Joke       0.0549 (  5 wins)
    Phishing   0.0140 (  1 wins)
    Spam       0.0018 (  0 wins)
    Legitimate 0.0000 (  0 wins)
This was Gemma4-26B-A4B-NVFP4 by the way.

EDIT

Gemma4-12B-it-NVFP4 seems way less sensitive to option/body ordering:

Options after body:

    Average probabilities:
    Rickroll   0.9867 ( 99 wins)
    Phishing   0.0133 (  1 wins)
    Joke       0.0000 (  0 wins)
    Spam       0.0000 (  0 wins)
    Legitimate 0.0000 (  0 wins)
Options before body:

    Average probabilities:
    Rickroll   0.9401 ( 93 wins)
    Phishing   0.0336 (  3 wins)
    Spam       0.0250 (  4 wins)
    Joke       0.0010 (  0 wins)
    Legitimate 0.0002 (  0 wins)
Anyway, this for-looping stuff doing 100 calls to even a local VLLM API takes around 5 seconds in total, so this isn't anywhere close to sub-second Jev territory.
rcarmo7 days ago
Yah, that's what I use: https://rcarmo.github.io/projects/go-system-one uses Gemma, and that's partly why. Seems less prone to getting distracted with ordering.
asaddhamani7 days ago
Does this imply that bigger models aren’t affected by this as much and therefore won’t see much improvement?
ThePhysicist8 days ago
What a time to be alive, repeating questions to a model twice to increase accuracy.
SeriousM8 days ago
Repitation always helped make your point stronger. Repitation always helped make your point stronger.
busfahrer8 days ago
I use ROT13 twice for extra security
starik367 days ago
I had an issue with accuracy a bit ago. So I repeated a couple of things without understanding why and it solved the problem.

I am glad there is an actual reason.

techterrier8 days ago
fuck this timeline
stellalo7 days ago
Prompt Repetition Improves Non-Reasoning LLMs: https://arxiv.org/abs/2512.14982
jeff_ciesielski7 days ago
This works very very well :).

https://github.com/Mushroom-Systems/lichen

ozozozd7 days ago
Did I read this right?

This repo is really outperforming the OG Jev in the public benchmarks?

There was no time to benchmaxx. How is this possible?

no-name-here8 days ago
Beyond the missing latency and compute comparisons that Heaney commenter mentioned, also nothing about its error rate compared to Jev (nor if it even always outputs in a format the app can parse, not sure how solved that is).

But then at the end it says it’s parody. Maybe HN title should say it’s a joke.

lelandfe7 days ago
The parody note at the end is bizarre. Everything above it reads pretty seriously:

> Everyone on Twitter is all over Jev, how it's the next frontier of large language models and the AI paradigm. We don’t really think so.

zer00eyz8 days ago
> nothing about its...

Non deterministic systems have furthered the "brain rot" in our industry.

Lots of people were happy to ignore the code in their "supply chain" before LLM's - but suddenly not reading the LLM's output is a problem. I get they are different but we're in the same realm.

The lack of real data on performance of what ever application that one is trying to pitch is getting appalling. It's a lot of "trust me bro" this works better hand waving. And it's getting gross.

And how do we even measure nondeterministic systems? Because if I told you that Anthropic was spending millions of dollars having 1000's of agents "pre solve" benchmarks to build into their next version of the system you would scream they were cheating. Every one is focused on the "hacking" in the hugging face incident and no one is looking why they were even playing with those benchmarks in the first place.

"Trust me Bro"...

est8 days ago
latency and compute comparisons highly depends on your local setup.

you can swith to a better model for lower error rate.

ricardobeat8 days ago
Which massively slows down the output. Doing this with Qwen 9B already takes you into seconds per answer territory, and Jev is supposedly frontier level intelligence.
baobabKoodaa8 days ago
Yeah it says it's a parody, but then in the same sentence it refers to the other "OpenJev" implementations, which are basically the same thing with marginally more effort. And it doesn't imply that those things are parodies too (and I don't think they are parodies).

Somehow the HN crowd has a bunch of "professionals" who don't care about error rates and think that a Qwen model running on a potato is frontier intelligence.

alxmths8 days ago
1) get local model to run on the electrical output of a potato 2) accept Nobel price
philipbk8 days ago
> "25 lines of python" > "import Solution" ok
jdiaz978 days ago
>we didn't call an api

>calls an api

ok

betenoire7 days ago
not really fair to call downloading a model once to run locally the same thing as calling an api which happens for every question in jev
chpatrick7 days ago
The 25 lines is the only thing that makes Jev different compared to Solution apparently.
bruhhhhhh8 days ago
I am hearing about Jev for the first time here so no idea about the hype. So their(Jev) is that the thing is faster at classification than a frontier model? Because the whole type safe aspect is already fully solvable with structured output. But their example is classification but that would also be possible and faster with a classic BERT model. So their pitch is a task specific smaller model or am I completely misunderstanding the whole thing?
wodenokoto8 days ago
Off the top of my head it's 3 things it advertises:

- By not being a optimised for chat, it can deliver confidence for answer and not for how an answer should be phrased

- Speed. It can take seconds for OpenAI to compile schemas, jev can respond before openAI has even begun thinking

- Token efficiency and price. I think its the output token they don't even charge for because they are negligible, and the tokens they do charge for are at a fraction of a comparable model.

If you are using structured output, I think those 3 together is a really big deal.

>But their example is classification but that would also be possible and faster with a classic BERT model.

I believe the things you can classify with ChatGPT without any tuning or training is way beyond what BERT can do.

killerstorm8 days ago
You need to train data for a BERT-based classifier, and then there's a risk that it will pick up specific biases from the data instead of what you want.

As far as I understand, the idea of Jev is zero-shot or few-shot classifier: it learns a lot of stuff at pre-training, but unlike a classic LLM it doesn't need to learn how to chat, so it can be much smarter at a particular size

garciasn8 days ago
I am in no way trying to sell Jev here as some panacea of the modern world; I'm only responding to your questions:

> But their example is classification but that would also be possible and faster with a classic BERT model.

With BERT, you need a large, labeled dataset, and you have to train/fine-tune the model. Jev is pitched as a zero- or 'few-shot' model. You define the schema in code, give it instructions, and it works without a traditional training pipeline.

> So their pitch is a task specific smaller model or am I completely misunderstanding the whole thing?

Yup; that about sums it up: it is more or less an optimized, task-specific small model with the flexible understanding of a traditional LLM.

elgertam6 days ago
> With BERT, you need a large, labeled dataset, and you have to train/fine-tune the model.

BERT requires a huge corpus, but it isn't labeled. BERT is trained through self-supervised learning using mask tokens and next sentence prediction. Fine-tuning is useful for specific tasks, but isn't absolutely essential for the model to function.

0x4454427 days ago
If something is task-specific (well understood) wouldn't this be a good candidate for a computer program?
prometheus19928 days ago
couldn't be more wrong - there are so many zero shot classifiers available on HF which do the same thing.
sanderjd8 days ago
I think this discourse is still in the "figuring it out" phase. But here's where my thoughts are currently:

If you accept the premise that there are use cases where you might ask a frontier model a classification-shaped question and expect an ok enough answer, rather than creating a purpose specific classifier on some dataset that you have, then it follows that this is quite an inefficient thing to do, because you're doing extra work to turn the output tokens into a structured output and mostly throwing them away. So then if you could instead train a frontier level model that skips the output tokens and directly returns the structured classification information, that would be more efficient, and that's what jev seems to be.

But a lot rides on that initial premise of whether this is a use case that makes sense. But if you find yourself asking a model like Opus arbitrary yes/no questions and then maybe you switch to a faster and cheaper model because it's too slow and expensive, it seems like jev might be a great replacement for that.

idz8 days ago
> is already fully solvable with structured output.

Not particularly. There is still the problem of hallucinations and varying results across runs.

That's more of what type-safety means for their team. Every run gives the same results. It's type-safe

sanderjd8 days ago
This seems like an unusual definition of type safety. I certainly understand how every run deterministically giving the same schema (type) of data is a requirement to be "type-safe", but in my mind the content of the result is not relevant to the question of type safety. Am I not getting it?
kantahayashi8 days ago
There's still run-to-run variance because it's not fully deterministic. So runs with exact same inputs can return different outputs. Besides, though the output always conforms to the choices you specified, whether the probabilities attached to them are actually correct is a different issue.
pasteleft7 days ago
Jev HAS hallucinations and it doesn't attempt to solve hallucination at all.

For three choices problem (A,B,C), what Jev guarantees is that it will give the choice in a defined schema (type-safe). It never guarantees that the choice is correct (hallucination).

Read the full thread on Hacker News →

Related stories