Fine-tunes of Qwen3.5 and Gemma 4 for zero-shot classification - firelex/jeff

570 points•firelex•2 days ago•223 comments•

223 comments

hadlock2 days ago
If you're looking for a more impressive doom example, I put together laya-duum which uses open source micropython implementation of doom (duum) and freeware freedoom1.wad. It uses the standard jev api and will play through the first two levels to completion: https://github.com/Hadlock/laya-duum
sunbum1 day ago
Why is this more impressive?
hadlock1 day ago
It's not just shooting at static monsters in an empty room, this will go after monsters, make decisions about prioritizing health, changing goals, and finish the level, etc. It's not unique, it's just more impressive than the demo they're using.
antman1 day ago
- Laya receives semantic snapshots, not the framebuffer.

What does tjis mean? Another generative model?

ovi2561 day ago
Laya is the model this uses. Like Jev, it's blind - it can't take an image as an input. So how can it play Doom, or Atari games, or do other visual tasks, as people have shown it to do? They write a bit of app-specific adapter code that transforms the game state into a structured piece of data (here, the "semantic snapshot"). And that is what these models take as input
prodigycorp1 day ago
laya sucks. it's a dead end and for some treat it as a peer of jev or even the qwen jevs. base model doesnt have enough world knowledge to generalize.
earino1 day ago
The base-model criticism is fair. Laya is weak zero-shot and doesn't have the world knowledge of a much larger pretrained model. Their own site puts the base models at around 0.35 on the typed-decisions benchmark, which is close to random.

I don't think that makes it a dead end. Laya is designed to be fine-tuned for a specific task, and the site reports the fine-tuning gains. The 0.766 number comes from fine-tuning on the benchmark's train split, not from the base checkpoint. They also report that fitting a single temperature scalar per question type cuts expected calibration error from 0.466 to 0.081. That's a large gain, and it only shows up after you specialize the model.

"Peer" is doing a lot of work in that comment. Jev can take on a new task without retraining because it starts with far more knowledge; Laya trades that away to stay small and trainable for a fixed task. So I wouldn't compare base Laya to Jev and stop there. Compare Jev to fine-tuned Laya on the same task and test set, then look at accuracy, latency, cost, calibration, and robustness, depending on which of those matter for the deployment.

vektormemory2 days ago
Can someone remove the extra LLM and just have an embedder do the classifier work?

It's turning into pimp my llm...

dingody1 day ago
I simply use an embedding model followed by a simple MLP, and that’s enough to solve many text classification problems.
2gay1 day ago
My little pony ?
nojs2 days ago
Wait until you hear about support vector machines!
abhgh2 days ago
Its funny - I was going to leave a similar comment - and I have, earlier, on a different thread. If people need fast classification, on a fairly scoped problem, it is very fruitful to start with an off-the-shelf embedding model like ModernBERT (which Laya uses) and stick a classifier in front - like a Support Vector Machine (SVM). For starters just tune the SVM, you don't even have to fine-tune the embedder - often it works very well, esp. given the compute needed. Plus you can get reliable confidence scores and generate explanations if you want them (using something like SHAP).
baobabKoodaa1 day ago
If you have an easy problem that can be solved by a classifier from pre LLM era, then sure, go ahead. But we have LLMs now and we can use those to expensively and slowly classify harder problems using general purpose models, without needing to spend a huge amount of time fine tuning a model for the specific task. Jev offers to do the same fast and cheap.

For clarity: no, a 0.8B model is not gonna do that.

trebligdivad2 days ago
What proportion of commercial LLM use is classification? I'm just wondering what happens to business AI spending/data centre usage when they realise they don't need full LLMs.
exogenousdata2 days ago
The American stock market probably loses 20% of its value in a few days.
drstewart1 day ago
Can you share your short positions so we can see how confident you are in these predictions?
causal1 day ago
Can you uhh elaborate on why?
svachalek2 days ago
I'd expect the vast majority at this point is coding. Classification is a thing but in my experience tends to run on light, cheap models, not the proprietary frontier ones.
RamblingCTO1 day ago
part of agentic engineering is classification as well tho. reviewing/gating, what to read etc. don't need a full LLM
BowBun2 days ago
We've stopped upgrading the models of our classification workflows for >1 year at this point, meaning they're running acceptably on early/mid-2025 models. That said I believe there is a long tail of non-production-ized users who throw this into their everyday LLM chats.
baobabKoodaa1 day ago
In one of my client projects we've had to do forced upgrades of LLM models because of deprecations (I believe three or four of them over the course of 2 years). Each time our internal benchmarks have shown REDUCED performance after upgrading to "better" models.
AgentMasterRace2 days ago
I compared it to Jev in my current use cases and it's very inaccurate. 70% vs 94% . for classification, it's unacceptable.
tbeseda2 days ago
For _your_ classification it's unacceptable. The OP seems to have anticipated this and mentions you can fine tune it for your use case. Did you try that?

I don't think the point is to displace Jev, but to show it's possible to build an MVP on open weights without years of work and millions of dollars.

Why (presumably) an engineer would dismiss exploring a lightweight, custom alternative to locking into a fashionable PaaS, I'll never know.

senko2 days ago
> The OP seems to have anticipated this and mentions you can fine tune it for your use case. Did you try that?

You can already so that with classification models such as ModernBERT, at 0.4B.

Jev's value is its zero shot performance without having to fine-tune.

nico2 days ago
Not sure the task at hand here. But if it doesn’t require any reasoning/thinking and it’s just a classification task, it’s worth a shot to look into training your own classifier

I’ve run some benchmarks. Using embeddings + logistic classifier, the architecture matches or beats Jev and Laya in all basic classification tasks (datasets tested: AG News, Emotion, MASSIVE Intent, Banking77) The type of task in which it does really well, especially against Laya, is classification with >50 classes

The classifiers also run in <1ms, so they can be very fast and precise at the same time

But this architecture has no “reasoning”, so it performs rather poorly on tasks that require it, like the ones from the XLNI dataset (Jev/Laya do a lot better on this one)

For the latter cases, you could use add a local lightweight LLM, something like a Gemma model. Or even some basic MLP, depending on the tasks/data

clhodapp2 days ago
Needing fine-tuning for the use-case completely changes the product category
baobabKoodaa1 day ago
You don't seem to understand what the point of Jev is when you say "you can fine tune it for your use case". Building your own classifiers for your specific business problems is the type of work we all used to do back in 2016 or so. It costs very much. Jev is a cheap and fast general purpose classifier.
Oras2 days ago
My use case is simple classification for job ads. Things like, industry, work settings (remote, hybrid, onsite) and job type (full time, part time .. etc).

I did side by side comparison with Gemini 2.5 Flash Lite, Jev, Jeff

I tried the 0.8B model, completely useless in classification. Qwen Jeff-Qwen3.5-2B was better, but still missed job type.

I suppose with larger model, this could be useful, but would require more ram and will be slower.

sharih1 day ago
most of these dev clones need fine tuning on your use case, might as well fine tune modernBERT then. Jev generalizes well while being fast and cheap
zergrush2 days ago
i've tried all the "open source" me too Jevs

they all suck

olwmc2 days ago
Sorry, do we have actual clear implementation details for Jev? I keep seeing these "recreations" or "Do Jev at home" but do we have access to their architecture? I haven't even used the product, I just find it strange.
Zetaphor2 days ago
What they did was immediately obvious and trivial to replicate
prodigycorp2 days ago
They make it clear that their edge is in their synthetic data. Train your own jevs miss the point. And if you have that much data, you didnt need jev or these replacements in the first place.

HN's desire to pretend the data pipeline doesn't exist or isnt meaningful is silly.

Read the full thread on Hacker News →

Related stories