We found an approach to get Jev-like properties from standard LLMs like GLM-5.3-Flash. The core idea is to craft the input prompt so that the first output token answers the question. This makes it possible to get a…

137 points•flxflx•4 days ago•59 comments•
We found an approach to get Jev-like properties from standard LLMs like GLM-5.3-Flash.

The core idea is to craft the input prompt so that the first output token answers the question. This makes it possible to get a decision with a single forward pass.

In the blog post, we describe the approach in detail for GLM-5.3-Flash and vLLM. We benchmark this setup against Jev and Laya. We find that our setup is on-par with Jev in terms of accuracy and speed and that it substantially outperforms Laya.

Still, in terms of costs per decision, Jev is several x better than our setup. In turn, our setup supports vision inputs.

59 comments

ricardobeat4 days ago
Everyone is doing this to emulate Jev, but...

I took a random book excerpt with 23,000 words (±30k input tokens) and used it as context. Jev still responds in 800ms, sometimes 500ms. That's in the neighbourhood of 20-50,000 tok/s prefill, which is obviously not possible with normal LLMs, not even Cerebras is this fast.

HarHarVeryFunny4 days ago
It seems people are just guessing at the architecture behind Jev. Obviously the functionality itself is easy to replicate, but why Jev seems to be making such a splash (beyond the doh! factor of it's huge applicability) is the ultra-low cost and speed, which may be due to architecture.

The Laya model compared in TFA shows one way Jev may be getting it's speed and low cost - by using a BERT-like bidirectional model rather than an auto-regressive one (LLM).

prometheus19924 days ago
was the answer correct?

i have tested jev for my use cases and its horrendously wrong, but then the follow up from jev's team is "oh, you need to boil the question down further". it's a spiral of how much do you wanna dumb down the ask so that it answers it correctly. i'll pass for now.

also, 30k input tokens is a lot.

amluto4 days ago
I imagine it’s not so hard to optimize a model for this use case.

Off the top of my head, I would skip all the modern linear attention / state space stuff and use classical attention. But run prefill in a fully sliding-window mode so that “state” tokens simply don’t attend to far away tokens, or maybe also allow everything to attend to the first few tokens (and train like this). Now prefill is almost embarrassingly parallel, and you can make it fully parallel by duplicating work at block boundaries. (I’m not saying this is an awesome architecture if you want excellent results, but I’m also not convinced that Jev gives excellent results…)

The let queries attend to everything.

And architect the stack around this. Don’t try to cache the KV data — process the queries as you go so that the each input block and layer’s K and V data is computed, attended to, and discarded.

I’m curious whether Cerebras actually is a good device for this. Cerebras is kind of low on RAM, but if you don’t need to store KV data, maybe the entire computation fits on the die.

mmastrac4 days ago
That's not true. I ran Cerebras as an experimental ultrafast Jev and it was faster.
ericpauley4 days ago
This has “/dev/null as a service” vibes…
walrus014 days ago
magic 8 ball as a service.
wongarsu4 days ago
I have run tests with qwen 3.8 and gemma 4 in a way similar to this post (based on an open source project that also does this with gemma4).

Getting competitive accuracy with Jev is fairly easy, if by accuracy you mean that the highest weighted answer is the right one. GLM 5.3 is complete overkill, much smaller llms will do

What Jev brings to the table, beyond speed, is that the reported probabilities match actual likelyhoods. If you present three options, with A and B equally likely and C impossible, jev will approximately answer with 0.5, 0.5, 0. Stock LLMs don't

walrus014 days ago
You can turn any sufficiently smart LLM into yes/no decision model or equivalent. I already have an existing workflow with a two paragraph detailed prompt, that sends pages of stuff to an LLM and asks it to return only 7 JSON objects. Several of those objects are binary "yes or no" choices of like, whether the content contains certain things.

You can even do it with small not particularly hard to host local LLMs like a variant of Qwen 3.6 35B A3B or 3.8 27B.

opiotrek4 days ago
But does it always 100% of the time sticks to the schema? We have a prompt that is explicitly instructed to return a single html tag with the response inside it and it sometimes hallucinates
gf0004 days ago
I'm fairly sure it is 100%, and has been available for "ages". That's how every tool call and whatnot works:

https://developers.openai.com/api/docs/guides/structured-out...

walrus014 days ago
yes, it does, with appropriate tuning/testing of the llm's parameters (temperature, top p, top k, using the right model, and the right prompt). You have to give it a rigid and very specific prompt to only answer in the JSON form. Also test it with LLMs that will handle being given a very low temperature to be very 'literal'. You're not asking for creative writing.

I should also add that the source comes from one of about 400 possible places and in a variety of messed up formats, it's the raw feed from a news scraper...

StevenWaterman3 days ago
Yes, you can use constrained decoding
m4y0u4 days ago
My question is why not use Jev instead? It's faster and cheaper.
kouteiheika4 days ago
> My question is why not use Jev instead? It's faster and cheaper.

Because it's proprietary? By using an open weight model you're guaranteed that you can access it forever; if one provider bans you then you can go to another one (or you can self-host). With a proprietary, single-provider model locked behind an API if your access is revoked you're screwed.

kylecazar4 days ago
There's some speculation that Jev is an open weight model with novel post-training (RLCD). So, if these folks have competitive accuracy with just the base model, it may raise some questions about the necessity of Jev's architecture. You generally don't want to find yourself competing only on price.

Fyi, I haven't tested this yet.

Fordec4 days ago
Also, while it's clearly got a lot of training on some use cases, others that probably weren't in the training set have worse good decision rates than a random number generator. If you can rebuild the architecture, you can train it on your use case.
HarHarVeryFunny3 days ago
Whether Jev is something architecturally different from an LLM (bidirectional vs auto-regressive? different type of parallel readout head?) remains to be seen, or guessed, but LLMs are certainly fungible and are competing on price - they leapfrog each other from release to release, but overall they are all progressing in unison.

When it comes to the high volume market of business automation, it seems that ultra-low cost, rather than expensive frontier intelligence, is exactly what you want, and low latency is also nice to have for customer-facing applications like customer service chatbots.

janalsncm4 days ago
> it may raise some questions about the necessity of Jev's architecture

When I hear “architecture” I am thinking number of parameters and latency.

When I hear “accuracy” I think training recipe, data, and (later) number of parameters.

So when you say that Jev’s architecture may not be necessary, the evidence I expect to see is comparable quality at comparable latency. Not equal quality at 2x latency and 4x the cost.

dcss_gardener4 days ago
I mean just from what's known of the funding and timeline it pretty much has to be based on open weights.

But it is likely more than just a fine tune + novel training. At the very least the LM head is swapped out for a classifier one and then or also idk, bidirectional attention for the encoding pass I'm out of my depth at this point and will stop guessing. The training is probably where they have the biggest moat though, not that it's necessarily huge.

I have a project that fits jev as advertised almost comically well and I've been playing with it, and the various hacks and open versions. Jev doesn't necessarily perform better overall but it is quite different. It's sensitive to prompt phrasing in ways the others aren't, it's easy to generate questions where all the other models cluster in confidence but jev is an outlier. Not necessarily more correct, but it does feel like it's getting its answers in a different way.

I'm guessing just as much as anyone else but I've been spending a ton of time on this the last couple weeks, it landed right when I was most ready to dig into it.

andrewchambers4 days ago
These questions are answered by the OP (Same speed, image support) - additionally, GLM is open weight.
schainks3 days ago
Compliance. You can’t put Jev in a HIPAA compliant service, for example.
vezycash3 days ago
Qwen3-Next-80B-A3B can already run on a 16GB M1 MacBook at around 3–5 tok/s using aggressive memory management. Could a Jev-style controller push that to 100 tok/s on an M1?

Read the full thread on Hacker News →

Related stories