Jev-like family of decision models built on top of Qwen3.5/3.8 you can train and run on your own - jaredpalmer/kev

462 points•tosh•10 days ago•205 comments•

205 comments

nico9 days ago
If you only need classification, and you can provide some training data, you can ask Codex/Claude to build an embeddings + logistic classifier model for you

For emails, I get 95% accuracy with this method, with only 50-100 examples for training

Training the model takes less than 5 minutes on a CPU

The resulting model is <1MB, and inference is sub 100ms

Some other cool things about this approach:

* the model doesn’t train on some “ideal” or general classification, instead it learns your preferences

* the model runs on pretty much any mobile device and can be retrained online on the device

* privacy, the whole training and inference is 100% local, no data goes anywhere (except whatever you feed codex/claude while building the model)

Note: to do a more general test, I made a classifier for the Banking77 dataset. The model is <10MB, trains in <30s on CPU and gets 94.5% accuracy, which puts it in the top 5?models by accuracy for that set (the best one is at 94.86%, but it’s 350MB in size and takes hours to train on a GPU).

psadri9 days ago
Now that the models are smarter / know more than most of us, a new bottleneck is finding out what you don't know. If you don't know about logistic classifiers and how they could be applied to your problem, you will not ask for it.

These days, I tend to start my coding sessions by the high level problem I'm trying to solve vs the prescriptive, specific solution I may have in mind. It often surfaces ideas and approaches that I did not know about.

dotancohen4 days ago

  > It often surfaces ideas and approaches that I did not know about.
I'd love to hear some examples. I do agree that tech is moving at a pace that I can no longer keep up, however that absolutely does not mean that I want a dependency on some project that didn't even exist this time last year.
0x4579 days ago
I built a whole thing that collects data, trains classifiers, exports models and dataset just for that. Claude writes me a terraform file that contains shape of the classifier and dataset. For images it can create datasets based of another dataset (crop this region from images that have these labels).

Originally it was so I can label data to fine-tune a VLM, but now a few tiny classifiers that run in milliseconds on cpu.

Now its collecting data to make a domain specific BERT and do what Jev does.

nico9 days ago
Very cool. What kinda of classifications are you running? How big are the models/training sets?

Also curious about if you plan on doing some sort of routing for the requests. Like detecting the type of task to decide which model to route the request to

swyx9 days ago
> Now its collecting data to make a domain specific BERT and do what Jev does.

this misses the point of jev somewhat - the point is that this is a foundational, general purpose classifier model - see some good sources https://x.com/mparakhin/status/2101683565520199887?s=12

constantlm9 days ago
I did this yesterday and it works incredibly well. I finetuned ModernBERT to classify documents. With zeroshot it achieved around ~30% accuracy, which jumped to 98.2% with finetuning, and latency of around 150ms on my Macbook. Just incredible!
roseway49 days ago
If you don’t see good performance with LRs, you may want to try RBF SVMs. We’ve found they work super well for our use cases with the embeddinggemma model as they can better separate classes in the non-linear embedding space.

Our resulting RBF models are tiny and fit in L1 cache, with microsecond inference latency.

tk909 days ago
I'm working on exactly this! I'm building a small model that classifies the correct DOM node containing an HTML's article content/title/date/author (given a raw html with a lot of noise/chrome). A fun learning exercise :)

30KB model, 40-50ms inference. Pretty happy with the results so far!

I can see an entire industry of tiny models like this, now that we have AI to help us do the grunt setup work (validation/training data creation, data cleaning, etc). Or just use a general classifier like Jev/Kev ha

Flere-Imsaho4 days ago
This. I see LLMs building smaller bespoke models on-the-fly to achieve a desired goal.
bicepjai9 days ago
Is this like for cleaning html data from say common crawl ?
ozozozd9 days ago
30kb model is super impressive.

What’s the model architecture?

prodigycorp9 days ago
Man, I'm already burnt out on all this jev talk.

The one thing jev has going for it is a dedicated company focused entirely on making the product good and keeping it maintained. I haven't been willing to jump on board with all these jev-shaped projects because their releases feel driven mostly by opportunism. I'm fine waiting a bit for the opportunists to shake out so we can see who is genuinely committed to bringing something valuable to the open-weight community.

Jev is much better than the traditional ML crowd gives it credit for, but my enthusiasm hits a wall when it comes to their data policy. It is completely draconian. Whatever you feed into the system, they retain.

The jev team needs to release a ZDR product, or their platform is dead on arrival. An open, jev-shaped model will win out solely on that basis.

magimas9 days ago
> I'm fine waiting a bit for the opportunists to shake out so we can see who is genuinely committed to bringing something valuable to the open-weight community.

that is generally a very healthy attitude in the AI space anyway in my opinion.

Some of our R&D departments haven't actually finished an interesting project in years because they keep jumping from trend to trend wanting to try out all the latest shit all the time.

jldugger9 days ago
Agentic has now replaced javascript for the "It has been 0 weeks since the last Y framework" meme.
kerwioru92384929 days ago
https://typesafe.ai/legal/privacy-policy

In their privacy policy they say

We (1) will not train or fine tune any artificial intelligence or machine learning models on Input, and (2) will not disclose any Input to a third party other than our service providers.

prodigycorp9 days ago
Yeah but they reserve the right to retain the data virtually indefinitely.

These aren’t acceptable terms on a personal or corporate level. I’ve seen some fools brag about proxying their life through jev. Messages, emails, LLM calls, files.

cle9 days ago
Will they train or fine tune on derivatives of input?

I'd prefer if these companies would just enumerate what they will do with my data rather than these vague over-specific claims about what they will not do, which leave me with more questions than answers.

preuceian9 days ago
What about throughput (reasoning) and output? Or can we reasonably assume those are downstream of input and therefore covered by this policy as well?
hhh9 days ago
They offer ZDR. It was effortless to get.
doublerabbit9 days ago
It's Openclaw all over again.
prodigycorp9 days ago
I was enthusiastic about the release of jev much more than I was openclaw because new generative primitive are fun to play around with. But this may be the fastest I’ve ever gotten to being sick and tired of the discussion cycle around it.
andriy_koval9 days ago
> The one thing jev has going for it is a dedicated company focused entirely on making the product good and keeping it maintained.

by not allowing to benchmark it, they make it user-hostile, you don't know for what kind of quality you pay

oscarfr10 days ago
Found this benchmark for Jev-class models: https://benchmarkheaven.com/jev-models

There are already many Jev-like models in there.

Edit: No affiliation. Just found it and thought others might find it interesting.

jasonjmcghee10 days ago
The open source ones- I downloaded a number and tried them and compared to Jev.

Anything that required knowledge / familiarity mmBERT and ModernBERT post-trains performed much worse.

So it seems like they did some kind of useful expansive pre-training.

Things that were Qwen or Gemma Diffusion did better at those kinds of tasks but were generally pretty inconsistent in terms of whether they could succeed repeatedly (and be stable + reliable) on the many types of tasks that are in the cookbook part of the Jev docs.

If you ask Jev similar input + questions, it's pretty stable. And does a reasonable job on a lot of questions.

This one public benchmark (the only I've seen) seems to give the open versions way too much credit. It wasn't my experience at all.

It gave my a false wrong sense of what might be required to get it working for something at work to avoid needing a new subprocessor as - at least on Cloudflare / OpenRouter Jev is third-party not hosted.

oscarfr9 days ago
Thanks for sharing your findings!

We are working on running our own benchmark of Jev and some of the other models. Our use case is classification that runs in a UI. Currently LLMs have good accuracy, but are too slow (and expensive).

Jev not being available through a cloud provider (Bedrock or similar) makes it more challenging for us to start testing and rolling it out.

raybb9 days ago
One thing I'm still trying to figure out is how this compares to something like gliner. If you're just doing classification in what situations would you choose kev vs gliner?
kingnetart9 days ago
totally, considering gliner2 (and especially 2.5) can do _most_ of things that jev can, minus the flashy demos
hbarka10 days ago
If Jev is fundamentally trained using RLCD while you’re building on a Qwen model that was trained using RLHF, how can the resulting model be considered Jev-like?
mohsen110 days ago
I can't find it but saw that if you give Jev English alphabet as choices and ask it in a loop what model it is, it would say Qwen

also tried myself: https://console.typesafe.ai/playground?share=shr_1690a3160f1...

prodigycorp10 days ago
People seem to turn their brain off when it comes to this type of cargo culting. This doesn’t mean much. Qwen often identifies itself as Claude. Does that make it Claude?
bityard10 days ago
That is not how models work.

Unless specifically told in a system prompt, the pile of weights has absolutely no knowledge of itself. You could hypothetically train it to answer such questions, but nobody bothers to do this, and ALL "knowledge" embedded in the weights is probabalistic anyway.

(I feel like this should be common knowledge in LLM discussions on HN by now.)

riedel9 days ago
Can someone explain to me how such self-awareness can be forced into the model. I mean I guess the pre training data could contain all sorts of stuff. How reliable are those hacks. I know that a lot of open weight models answer that they are Claude in the absence of a system prompt. I find destillation not that much of a plausible explanation as typically claude would probably not mention that it is Claude all the time. I find it rather plausible that a foreig. system prompt made it into pre-training. But again: I have no clue how much care is given by models to leave traces for destillation (for closed weights) or post training (for open weights).
fxwin10 days ago
for some reason this is really funny to me. it's like the "black museum" black mirror episode where a consciousness in a toy animal can only communicate using very primitive predefined responses
c7b10 days ago
Once the first letter is Q, the rest is probably pretty determined. Can you see the confidence for the first letter (don't want to accept the ToS to follow your link)?
llm_nerd10 days ago
> If Jev is fundamentally trained using RLCD

Big if. More likely, it seems, is they started with an open LLM model and fine-tuned and repurposed it via their "RLCD" process.

jrmg10 days ago
The FAQ on the Jet announcement (https://typesafe.ai/blog/introducing-system-one-models-and-j...) claims it was trained with their own data:

Where Does Our Training Data Come From?

TypeSafe is primarily a data research lab, which is how the biggest results in AI get made. We make all the data ourselves. We wouldn’t train on your data even if you asked us to (no offense). We do some pretty sophisticated stuff, but if you want to find out more, we’d have to hire you.

tietjens10 days ago
Also my question.
monkeydust10 days ago
Bit of a Jev explosion going on. Is it because it's taking us back to a simpler time we understand better? Classification models have been around for a while.
reacharavindh10 days ago
The way I see this (I havent played around with Jev or layla the OSS version) is that classifiers have always existed and a recognised tool in the ML world. But, the norm is that one needs to not only know what to classify as, but determine what weights to use to classify the input.

Jev came in, and added that magic of "you dont need to train your classifier or determine the weights" if you dont want to, and just get the classified answer out. I think that's what is making people see this with a glitter in their eyes.

justincormack10 days ago
I would be curious to see comparisons of jev and similar things with problem specific classifiers. I think layla suggested making problem specific versions anyway? There is a lot of demand for magic don't do any work solutions, which is kind of weird in an era where agents can really help you build a customised solution effectively.
jsw9710 days ago
Agreed.

Just to be helpful if anyone is searching for layla, it's laya.

nater50009 days ago
>Classification models have been around for a while.

I'm still trying to catch-up on the Jev stuff, but my understanding is that it's basically just a more efficient LLM when all you want is the LLM to produce a classification.

There's more to it, of course, but it's not just "generic" classification ML because it accepts arbitrary inputs and can produce probabilities over arbitrary classes. Not saying this is the first time people have done this, but typically classification tasks are more static and limited.

In the same vein, it's also not just an LLM with structured outputs (which have been a thing for a while) specifically because that is a very inefficient way to approach classification using this kind of architecture. Jev models are much more performant because of how limited they are compared to a full LLM.

So when you want an LLM, but you only really need this kind of classification from the LLM, then Jev makes a ton of sense. This makes sense for me, since I've definitely used LLMs for this kind of classification work and, even then, it kind of felt like using a jackhammer to place some nails, etc.

Happy to be correct, though.

ryeights9 days ago
But an LLM provider could very easily add a "Jev mode" to any existing model, right? LLMs already produce a probability distribution over arbitrary classes. Just tell e.g. 5.6 Luna “here is the user's question, you must respond ONLY with the words 'foo', 'bar', or 'baz',” run a single forward pass of the model, and report the normalized probabilities of 'foo' 'bar' and 'baz' tokens before the first output.

With such an approach you could even retain full reasoning capability

dwedge9 days ago
> Happy to be correct, though.

Not normally one to point out a typo but this one made me smile

badatnames10 days ago
It reminds me a bit of what Ansible got right: user communication. The underlying tech may have existed for a long time, but the genius is presenting it to a regular developer in a way that reads "yes, even you can understand ML, just using a little JSON". The contribution of that should not be understated, as has been clearly evident recently.
colordrops10 days ago
Yeah except it doesn't really work. It constantly breaks underneath you. The whole system has to be managed, e.g NixOS, or else it's a house of cards.
whazor10 days ago
The Jev model is economically, but also in terms of compute, a much more efficient model. A normal LLM goes token by token, each token in a separate step. Whereas Jev just returns all the results the first round. So it is much better at classification than LLMs.

Compared to traditional ML classification, Jev works without training, like a LLM.

qudat10 days ago
Previous classification models need to be trained on the specific question/choices you are trying to output. Jev doesn't need to be retrained for every choice set provided.

LLMs can act as classifiers but they still have to generate text output in the form of a JSON object. This means they have to generate every single curly bracket, quote, command, etc. This turns out to be pretty expensive. On the other hand, Jev uses a different decision head so it doesn't generate text output at all, it outputs logits *only* for the choices provided. So it completely avoids the need to generate text at all, which means no malformed JSON and it's much faster as a result.

Finally, Jev also provides confidence scores that are actually reliable (not made up like LLMs).

Read the full thread on Hacker News →

Related stories