Compare Jev-class decision models across intelligence, calibration, speed, and cost with JevBench.

153 points•florianstandhar•9 days ago•39 comments•
Hi HN! I built JevBench because Jev kicks ass, and the world deserves to know how the serious open source and fake lookalike projects really perform in comparison.

Jev-class models return bounded choices and probabilities instead of text, and are disruptively faster and cheaper than LLMs, while being similarly intelligent on the text input they operate on.

JevBench allows looking at accuracy, latency and price all at once, in a weighted way - you can even configure the weighting.

A full run asks 534 English decisions. The v1.3 score combines chance-corrected Intelligence, Calibration, Speed and Cost.

Leaderboard right now:

  #1 - Jev            74.4
  #2 - SemIf          73.1
  #3 - djev           73.0
  #4 - Winnow-12B Q8  71.2
  #5 reflex 4B        70.3.
MIT harness, public items, frozen artifacts, scoring code and public per-task outcomes:

https://github.com/fstandhartinger/jevbench

Two no-signup demos:

https://who-is-right.app.mintapis.com

https://is-it-ai-slop.app.mintapis.com

Limitations: English-only; latency from one German server; local/demo latency gets a disclosed ×2 adjustment (+150 ms on my servers) which is an informed assumption; held-out prompts still reach evaluated services; ~1-point gaps can be noise.

Wdyt?

39 comments

hbrn8 days ago
$40m in funding, 2 years in stealth.

Performs on-par with SemIf which was built in a couple days and apparently uses raw Qwen, with no fine-tuning. SemIf runs in your freaking browser. Oh and Jev is twice as expensive?

Is it surprising that Jev consistently thinks it's Qwen?

I'm almost convinced that Jev is a scam. Take Qwen, fine tune it a little, tell investors it cost $10m, spend $1m on advertising, profit.

vessenes8 days ago
I think it’s much more likely that these benchmarks are not good, in that they do not explore much of the space that jev was (likely) trained to cover. The doom demo is a good example — I’d like to see a wide variety of things like that included in any benchmark, not just ‘email classification’ or what have you.

Think of it this way: there’s some time needed to optimize / design an architecture, and the world gets that for free when it’s described. As to the rest of the last two years spent, is it more likely a former oAI lead spent them fucking around, or adding as many RL environments as possible to its model that is supposed to be a generalized classifier?

Right now my prior is that jev is probably better than these rando weekend models, whether or not we know how to test and demonstrate that in a benchmark. It’s also super cheap, so I don’t think there’s a strong reason not to try out building with it first, then walk down the ladder to an open model if you need to for some reason.

hbrn7 days ago
So if benchmark says Jev sucks, it's benchmark's fault. And yet Typesafe doesn't share anything about their internal benchmark that apparently proves how amazing Jev is.

Funny enough, when I tell people that Jev cannot play tic-tac-toe I hear a similar argument - it's not what Jev was built for. Jev has this elusive use case that noone can describe, so when Jev fails, it's just because it wasn't built for it. Convenient.

Don't you find it suspicious that Jev cannot play tic-tac-toe or checkers, but can play Doom? Don't you find it suspicious that nothing of the Doom demo was shared: no harness, no control loop, no state encoding, no prompts - nothing.

And the explanation is obvious - Jev isn't playing doom. The harness is. They essentially built a Doom bot, dumbed down it's control loop for Jev, and gave reins to Jev. Look mom, Jev is playing Doom!

The state says "you're pointing at the cacodemon" or "you're not pointing at the cacodemon" and Jev has to decide whether to press fire. Frontier intelligence!

You can build a harness where a coin flip is playing Doom.

florianstandhar8 days ago
Agree the coverage question is real. Today it's text decisions (intent, routing, extraction, scoring, harder multi-constraint cases). We're working on screen/computer-use and image decisions (a noindex preview exists) because that's closer to where Jev-class models will actually be used. Suggestions for task families are very welcome as GitHub issues.
Retro_Dev8 days ago
Isn't the entire deal with jev that it is fast? I'd be interested to know how the energy cost of the Qwen-based model compares with Jev. Of course, Jev is currently locked up so we don't know... "Trust me bro Jev is revolutionary and amazing, pay more money for our inferior product which costs more to run, and of which you need to access by sending us the data"
dmix8 days ago
You can spot vibecoded websites by how they include the prompt or commit-style comments into the literal interface, instead of communicating it via visual context (or simply excluding it)

> Sort by any column; values the run could not produce always sort last. Hover a cost for how it was priced, a latency for the endpoint. Names link to each project.

A designer would never write this, but an LLM just inserts it by making it small grey text next to the interface, just like it does with inane code comments.

spiderfarmer8 days ago
EYEBROWS EVERYWHERE

There are lots of tells in every vibe coded design.

anukin8 days ago
And why is a designer needed for every website in the world?
wannabe448 days ago
I (not GP) have no issues with CSS part of the design (just as I am fine with bootstrap websites). But if you can't be arsed to write 4 lines of text describing it in human language instead of the claude vomit, you probably didn't put much effort to begin with.
spiderfarmer8 days ago
They're not needed. And there are bad designers everywhere.

And I know HN people think marketing is a bad word.

But marketeers know that design is part of the messaging. And nowadays, having clear AI tells in your design shouts "I made this in an afternoon, so don't take this seriously"

First impressions are everything in a world where attention spans are shortened every year.

Act accordingly.

dogscatstrees8 days ago
It's helpful metadata, what's wrong with it? Dashboards at work do this.
janalsncm8 days ago
One person’s helpful metadata is another person’s noise. It’s much better for a door to visually indicate that it should be pushed open than to put up a sign there.

However, doing the former requires a level of empathy with humans that LLMs rarely have.

Human brains have caloric demands. It is possible for humans to process enormous amounts of unrelated facts to make a decision, but it’s tiring. It’s much better to not do that, especially just to get some basic information.

To anthropomorphize a bit, an LLM might find it charming and interesting to read someone’s life story as a preamble before their taco recipe. Humans by and large find that annoying, not because we can’t understand the biography but because processing that information is not free.

So it’s probably possible to design using an LLM. You would probably have to be intentional about it.

CBLT8 days ago
If you don't have any kind of agent instructions saying "don't copy code into prose", you'll inevitably get text that reflects some previous state of the system instead of its current state.
dmix8 days ago
It's fine in isolation to have help text for complex interfaces needing explanation. Better yet contextual hints.

But there's about 10 other examples explaining intention of the coding rather than immediately useful information to the user, it's all over this one site in small grey text. And I guarantee you nobody is reading them carefully. Just like how nobody likes reading a 15 line LLM code comment over a simple function.

Most of it could be better solved with more thoughtful design or deleted. The link explanation is particularly egregious.

sean_pedersen8 days ago
Good project but this one also exists https://huggingface.co/spaces/multimodalart/jev-decision-ind... and the results do not seem to add up and also model sets are different... still needs time to mature likely
ks20488 days ago
I was trying to figure out what exactly the tests here are. I guess I found some of the questions (here: https://github.com/fstandhartinger/jevbench/blob/main/datase...)

e.g.,

  "instructions": "Which intent does the user's message express?",
  "labels":["set_alarm", "play_music", "weather", "send_message", "turn_off_lights"],
  "state": "Play some Taylor Swift.",
  "expected": "play_music"
swyx8 days ago
arbot3608 days ago
Many SaaS vendors forbid benchmarking, I find it crazy that such anti-competitive terms are standard across the industry but they are. Generally the goal of such terms is to "control the narrative" around the product, regardless of the truth of performance being better or worse than competitors.
meander_water8 days ago
They released some examples of what their workflow evals are like. I'm sure you could reverse engineer a benchmark from that

https://evals.typesafe.ai/

jldugger8 days ago
Interesting; was curious how this didn't fall into trouble with ToS. Apparently the "no benchmarks" clause was intended for "limited preview" audiences and didn't get removed at launch on accident.
Maxious8 days ago
But also no big deal because "I’m extremely anti-public benchmarks."
tomrod8 days ago
I mean, that's a great reason to ignore JEV entirely.

"Trust, but verify" isn't just a catchy cliche. It's the only way to operate where models and code are fast to market.

Read the full thread on Hacker News →

Related stories