Jev-like family of decision models built on top of Qwen3.5/3.8 you can train and run on your own - jaredpalmer/kev
205 comments
For emails, I get 95% accuracy with this method, with only 50-100 examples for training
Training the model takes less than 5 minutes on a CPU
The resulting model is <1MB, and inference is sub 100ms
Some other cool things about this approach:
* the model doesn’t train on some “ideal” or general classification, instead it learns your preferences
* the model runs on pretty much any mobile device and can be retrained online on the device
* privacy, the whole training and inference is 100% local, no data goes anywhere (except whatever you feed codex/claude while building the model)
Note: to do a more general test, I made a classifier for the Banking77 dataset. The model is <10MB, trains in <30s on CPU and gets 94.5% accuracy, which puts it in the top 5?models by accuracy for that set (the best one is at 94.86%, but it’s 350MB in size and takes hours to train on a GPU).
These days, I tend to start my coding sessions by the high level problem I'm trying to solve vs the prescriptive, specific solution I may have in mind. It often surfaces ideas and approaches that I did not know about.
> It often surfaces ideas and approaches that I did not know about.
I'd love to hear some examples. I do agree that tech is moving at a pace that I can no longer keep up, however that absolutely does not mean that I want a dependency on some project that didn't even exist this time last year.Originally it was so I can label data to fine-tune a VLM, but now a few tiny classifiers that run in milliseconds on cpu.
Now its collecting data to make a domain specific BERT and do what Jev does.
Also curious about if you plan on doing some sort of routing for the requests. Like detecting the type of task to decide which model to route the request to
this misses the point of jev somewhat - the point is that this is a foundational, general purpose classifier model - see some good sources https://x.com/mparakhin/status/2101683565520199887?s=12
Our resulting RBF models are tiny and fit in L1 cache, with microsecond inference latency.
30KB model, 40-50ms inference. Pretty happy with the results so far!
I can see an entire industry of tiny models like this, now that we have AI to help us do the grunt setup work (validation/training data creation, data cleaning, etc). Or just use a general classifier like Jev/Kev ha
What’s the model architecture?
The one thing jev has going for it is a dedicated company focused entirely on making the product good and keeping it maintained. I haven't been willing to jump on board with all these jev-shaped projects because their releases feel driven mostly by opportunism. I'm fine waiting a bit for the opportunists to shake out so we can see who is genuinely committed to bringing something valuable to the open-weight community.
Jev is much better than the traditional ML crowd gives it credit for, but my enthusiasm hits a wall when it comes to their data policy. It is completely draconian. Whatever you feed into the system, they retain.
The jev team needs to release a ZDR product, or their platform is dead on arrival. An open, jev-shaped model will win out solely on that basis.
that is generally a very healthy attitude in the AI space anyway in my opinion.
Some of our R&D departments haven't actually finished an interesting project in years because they keep jumping from trend to trend wanting to try out all the latest shit all the time.
In their privacy policy they say
We (1) will not train or fine tune any artificial intelligence or machine learning models on Input, and (2) will not disclose any Input to a third party other than our service providers.
These aren’t acceptable terms on a personal or corporate level. I’ve seen some fools brag about proxying their life through jev. Messages, emails, LLM calls, files.
I'd prefer if these companies would just enumerate what they will do with my data rather than these vague over-specific claims about what they will not do, which leave me with more questions than answers.
by not allowing to benchmark it, they make it user-hostile, you don't know for what kind of quality you pay
There are already many Jev-like models in there.
Edit: No affiliation. Just found it and thought others might find it interesting.
Anything that required knowledge / familiarity mmBERT and ModernBERT post-trains performed much worse.
So it seems like they did some kind of useful expansive pre-training.
Things that were Qwen or Gemma Diffusion did better at those kinds of tasks but were generally pretty inconsistent in terms of whether they could succeed repeatedly (and be stable + reliable) on the many types of tasks that are in the cookbook part of the Jev docs.
If you ask Jev similar input + questions, it's pretty stable. And does a reasonable job on a lot of questions.
This one public benchmark (the only I've seen) seems to give the open versions way too much credit. It wasn't my experience at all.
It gave my a false wrong sense of what might be required to get it working for something at work to avoid needing a new subprocessor as - at least on Cloudflare / OpenRouter Jev is third-party not hosted.
We are working on running our own benchmark of Jev and some of the other models. Our use case is classification that runs in a UI. Currently LLMs have good accuracy, but are too slow (and expensive).
Jev not being available through a cloud provider (Bedrock or similar) makes it more challenging for us to start testing and rolling it out.
also tried myself: https://console.typesafe.ai/playground?share=shr_1690a3160f1...
Unless specifically told in a system prompt, the pile of weights has absolutely no knowledge of itself. You could hypothetically train it to answer such questions, but nobody bothers to do this, and ALL "knowledge" embedded in the weights is probabalistic anyway.
(I feel like this should be common knowledge in LLM discussions on HN by now.)
Big if. More likely, it seems, is they started with an open LLM model and fine-tuned and repurposed it via their "RLCD" process.
Where Does Our Training Data Come From?
TypeSafe is primarily a data research lab, which is how the biggest results in AI get made. We make all the data ourselves. We wouldn’t train on your data even if you asked us to (no offense). We do some pretty sophisticated stuff, but if you want to find out more, we’d have to hire you.
Jev came in, and added that magic of "you dont need to train your classifier or determine the weights" if you dont want to, and just get the classified answer out. I think that's what is making people see this with a glitter in their eyes.
Just to be helpful if anyone is searching for layla, it's laya.
I'm still trying to catch-up on the Jev stuff, but my understanding is that it's basically just a more efficient LLM when all you want is the LLM to produce a classification.
There's more to it, of course, but it's not just "generic" classification ML because it accepts arbitrary inputs and can produce probabilities over arbitrary classes. Not saying this is the first time people have done this, but typically classification tasks are more static and limited.
In the same vein, it's also not just an LLM with structured outputs (which have been a thing for a while) specifically because that is a very inefficient way to approach classification using this kind of architecture. Jev models are much more performant because of how limited they are compared to a full LLM.
So when you want an LLM, but you only really need this kind of classification from the LLM, then Jev makes a ton of sense. This makes sense for me, since I've definitely used LLMs for this kind of classification work and, even then, it kind of felt like using a jackhammer to place some nails, etc.
Happy to be correct, though.
With such an approach you could even retain full reasoning capability
Not normally one to point out a typo but this one made me smile
Compared to traditional ML classification, Jev works without training, like a LLM.
LLMs can act as classifiers but they still have to generate text output in the form of a JSON object. This means they have to generate every single curly bracket, quote, command, etc. This turns out to be pretty expensive. On the other hand, Jev uses a different decision head so it doesn't generate text output at all, it outputs logits *only* for the choices provided. So it completely avoids the need to generate text at all, which means no malformed JSON and it's much faster as a result.
Finally, Jev also provides confidence scores that are actually reliable (not made up like LLMs).
Read the full thread on Hacker News →
Related stories
- The Verge · 0 points · 5 days ago
- The Verge · 0 points · 3 days ago
- Can John Ternus find Apple’s next big thing?theverge.comThe Verge · 0 points · 10 days ago
- Lobsters · 86 points · about 1 year ago
- The Verge · 0 points · 12 days ago
- Can you forget how you feel about Meta?theverge.comThe Verge · 0 points · 9 days ago