TypeSafe's Jev is a genuine breakthrough – snap judgments with calibrated probabilities instead of generated text. My bet is OpenAI is already figuring out how to copy it, and then embed it inside its own models where…

327 points•JohnBerryman•9 days ago•229 comments•

229 comments

orbital-decay8 days ago
Every major AI shop has a ton of in-house classifiers already, big, small, generalist, specialized. Some are used in inference pipelines (e.g. safeguards), some are used in data preparation, training, analysis and investigation, research, various one-off and intermediate tasks etc. Offering them on a public API doesn't always make business sense. I don't see much substance to this buzz, looks like people that are new to all this are discovering that classifiers exist, they are more efficient at classification, and many tasks commonly done with generative models are classification in disguise. Which is not bad at all, a fresh look at their use is great to have.
bigmadshoe8 days ago
Correct me if I'm wrong, but a zero-shot classifier like Jev is fundamentally different to a classifier with a fixed task (e.g. for safeguards), unless they trained a general purpose system to complete the safeguard task, which seems unlikely.
zxexz8 days ago
This whole thing reminds me of DeepMind’s Variational Bayesian Last Layers[0], which never gained much traction in the broader “AI” world, but is a remarkably useful tool. And a relatively obvious one that anyone with experience in SVI and with transformer pretraining, seems to independently rediscover (including me) before finding this paper.

[0] https://arxiv.org/pdf/2404.11599

janalsncm8 days ago
Correct, but zero-shot classifiers are also not new.
BiteCode_dev8 days ago
The generalist aspect of jev is what is good about it. Old classifiers tended to be specialized, not good at ambiquity or limited.

LLM were used as classifiers because they solved that.

Jev have the flexibility of LLM and the perf and api of classifiers.

mmis10008 days ago
Fixed guard today is not very fixed. For ex, the safeguard qwen released is a full 4b llm model. It has no different to normal llm model arch except tuned for this specific purpose,
EagnaIonat8 days ago
I fed into the hype at first. Testing Jev and Laya, they both suffer from the same issues as LLMs that stop them being useful beyond limited classifications.

I can't see any benefits that a typical ML classifier would not be better at.

edot8 days ago
Agreed. I tested Jev on OpenRouter this past weekend and it’s “okay” but a specific classifier is significantly better. It used to require skill to import sklearn (ok, not really), but now it’s literally one prompt and upload your Excel file or whatever and you can get your classifier out. It’ll run free, instant, more accurate.
ainch8 days ago
I think the main argument would just be that because the model is general, you don't need to retrain it from scratch for a new problem - just tweak the input prompt. For a typical classifier there's a lot more hassle - collecting the data, training it yourself, retraining under distribution shift... In that sense Jev seems great for prototyping or small-scale use cases.
ricardobeat8 days ago
Using Jev as a plain classifier is the least interesting case. See robotic control, navigation, computer use examples, none of it possible with a classifier.
tomrod8 days ago
Prompt ingestion is going to be the biggest differentiator.

Being able to route prompt to features that then route to special models would be a really solid implementation.

sanderjd8 days ago
I guess I'm circling toward this view. The question is, are there things that are 1. worth doing, 2. for which jev (or jev-like systems) works well, and 3. are not worth the effort to train a custom classifier. Probably yes, but it seems like it might be a pretty narrow path. But a lot depends on #2. The trade-off between #1 and #3 is less stark the more successful one shot models are at handling use cases successfully.
d2ou8 days ago
People in my lab (sklearn people) developped something that I feel close to jev but focused on tabular data : https://tabicl.readthedocs.io/en/latest/

This is a transformer based classifier with massive pretraining on synthetic datasets and it outperforms boosting classifiers on many benchmarks without the need of more gradient descent steps (the forward pass on X_train, y_train IS the training).

I understand that jev focus on text entry. But I feel that it is a similar kind of model but trained on text. Did someone test it on tabular data as well ?

gwern8 days ago
Entertainingly, OpenAI had a general purpose zero-shot classifier API built on GPT-3! Just no one ever cared that much about it, so I guess it got dropped somewhere along the way since 2020/2021.
0x20cowboy8 days ago
This. It’s machine learning vs. “AI” for the uninitiated. Soon there will be a new ground breaking model that does k-means clustering and will get a billon dollar funding (but only if you're young and live in SF)

The good news is it’s fun to see people discover and get excited about things that I like as well.

aDyslecticCrow8 days ago
Its in a modern and easily to deploy package. The hype is a bit wierd. "0 cost output tolkens" is such a silly phrasing.

I would have never considered importing pytorch for filtering through log files before even knowing my way around it. But if i can type a filtering condition by text and hit enter; i may actually use that to save some time.

Id want something local though, but thats hardly a difficult demand for what it is.

prodigycorp9 days ago
This article is extraordinarily hard to read. It’s tummelvisioned on OpenAI and things like tool calling which are only relevant to the extent that llms have been tuned to make relative choices, but this applies to all LLMs. Also, some really outdated references. LLM written, perhaps?

Also, moat discussion is the lowest form of discussion. I don’t care if jev has a moat. Did it get the interface right? What other past ideas have we overlooked that if given some love, could kick the door down like jev did?

Really silly stuff.. people wanting to talk about moats when there’s no castle. Moat talk merely projects the illusion of being engaged but, much more often than not, it’s hollow engagement.

jrochkind18 days ago
> LLM written, perhaps?

pangram says... 20% of content likely AI written, 80% of content likely human written.

Eventually humans are going to start writing like AI if we read enough of it.

JohnBerryman8 days ago
Partially. I'm a terribly slow writer and get stuck on phrase choice, but I'm good at content ideas, outlines, and editing text already on the page. So I have AI do the bit that I'm not as good at.

Process: First, actually have ideas :D Then, I write an outline for what I want to talk about at basically a sentence-by-sentence level. (This is me yelling things at my computer.) And then I have the AI convert a chunk at a time into prose. I reread it and rework it to be my voice.

Then I have the AI help with things like subject titles and social posts.

¯\_(ツ)_/¯

andy12_9 days ago
I find it unlikely. OpenAI is all in training models with reasoning with RL, and Jev-like models are the total opposite. They are made to not reason at all to be fast. If you want to add reasoning on top, you might as well use a conventional LLM because you lose the price and speed benefits when you output auto-regressive tokens. I don't think OpenAI will even bother with this.

> My main assumption is that Jev is using something quite close to a conventional large language model. As evidence of this, Latent Space reports that many of the early clones are indeed LLM-based.

Not proof that this is the case with Jev though. It might use non causal text encoder for the state, which could make sense given that it's very good for its price.

altcognito9 days ago
I don't see it fundamentally any different than knowing when to use a tool. Is this tool like RAG an important enough corner case to train for it? I dunno.

LLMs already shell out and write code to solve certain problems. This is just a special case of that.

andy12_9 days ago
It's a special case for an LLM, and you can use an LLM with structure output to get similar results, but you can engineer specifically for that case to get better results per dollar for it. That's why there is little reason to adapt GPT 5.6 Sol or wathever for this task; it can already do it (at a high cost). For OpenAI to compete with Jev they have to maintain another line of models, something like "GPT-5.6-instant-decision", that is small, fast and cheap, in the scale of GPT-5 nano.

Note that I don't think OpenAI is incapable of doing it, but I just don't think they will bother with it.

Onavo8 days ago
In the olden days we call this classifier, usually assignment 2 of Machine Learning 101. BERT (well, GLiNER specifically) and diffusion are calling and want their Large Classifier Models back.

https://github.com/vllm-project/vllm/pull/57250

scottyah8 days ago
If there's money to be made, I'm sure sama will find a righteous reason to offer it.
himata41139 days ago
system 2 is just an llm with a forced toolcall IMO
rdevsrex9 days ago
There is one benefit that Jev has, that it is not OpenAI and thus it's probably less likely to steal your own work.
docheinestages9 days ago
> that it is not OpenAI

For now. Any company that grows to OpenAI/Anthropic's size and gets VC money is ought to become greedy.

Andrex8 days ago
Or OpenAI just buys them outright. Buying your upstart competitor seems to be in the Silicon Valley Ten Commandments. The Fed whussed-out on breaking up FB and Insta last year, so there's never going to be any kind of remediation to worry about.

And for Jev, everyone has a price, and OpenAI's raised an historical amount of funding.

SubiculumCode9 days ago
Not steal, but keep it indefinitely, per the JEV TOS
Synthetic73469 days ago
They absolutely need ZDR
linhvn8 days ago
Every company would use your data and "steal" it at some point. VC money would demand it the moment the pressure is up.
heaney-5559 days ago
This is a tired argument that needs actual evidence to go beyond the level of a conspiracy theory.
Tanjreeve8 days ago
TIL the terms of service are a conspiracy theory.
zergrush9 days ago
comments are pretty weird here, there's no real moat to what jev is doing, it is certain that frontier labs are going to release their own jev and there are even open source alternatives (although nowhere near as accurate as jev).

so maybe typesafe's real plan is to front run and releasing their own new models for some time until they can get acquired which seems to be the only rational objective

jackb40409 days ago
I don't think it's unreasonable to think this, but I do think the burden of proof is on your side. Between the SaaS-pocalypse narrative that never materialized, and inexplicably losing their first-mover advantage to Anthropic, OpenAI's track record is not great when it comes to jumping on these micro paradigm shifts.

If the headline said "Frontier labs are about to eat Jev's lunch" it might be an easier sell. But if we're gonna include Anthropic, I think part of their success is actually making products for which there is demand. It will take time for something like that to come out of this new "decision model" paradigm.

Read the full thread on Hacker News →

Related stories