Fine-tunes of Qwen3.5 and Gemma 4 for zero-shot classification - firelex/jeff
223 comments
What does tjis mean? Another generative model?
I don't think that makes it a dead end. Laya is designed to be fine-tuned for a specific task, and the site reports the fine-tuning gains. The 0.766 number comes from fine-tuning on the benchmark's train split, not from the base checkpoint. They also report that fitting a single temperature scalar per question type cuts expected calibration error from 0.466 to 0.081. That's a large gain, and it only shows up after you specialize the model.
"Peer" is doing a lot of work in that comment. Jev can take on a new task without retraining because it starts with far more knowledge; Laya trades that away to stay small and trainable for a fixed task. So I wouldn't compare base Laya to Jev and stop there. Compare Jev to fine-tuned Laya on the same task and test set, then look at accuracy, latency, cost, calibration, and robustness, depending on which of those matter for the deployment.
It's turning into pimp my llm...
For clarity: no, a 0.8B model is not gonna do that.
I don't think the point is to displace Jev, but to show it's possible to build an MVP on open weights without years of work and millions of dollars.
Why (presumably) an engineer would dismiss exploring a lightweight, custom alternative to locking into a fashionable PaaS, I'll never know.
You can already so that with classification models such as ModernBERT, at 0.4B.
Jev's value is its zero shot performance without having to fine-tune.
I’ve run some benchmarks. Using embeddings + logistic classifier, the architecture matches or beats Jev and Laya in all basic classification tasks (datasets tested: AG News, Emotion, MASSIVE Intent, Banking77) The type of task in which it does really well, especially against Laya, is classification with >50 classes
The classifiers also run in <1ms, so they can be very fast and precise at the same time
But this architecture has no “reasoning”, so it performs rather poorly on tasks that require it, like the ones from the XLNI dataset (Jev/Laya do a lot better on this one)
For the latter cases, you could use add a local lightweight LLM, something like a Gemma model. Or even some basic MLP, depending on the tasks/data
I did side by side comparison with Gemini 2.5 Flash Lite, Jev, Jeff
I tried the 0.8B model, completely useless in classification. Qwen Jeff-Qwen3.5-2B was better, but still missed job type.
I suppose with larger model, this could be useful, but would require more ram and will be slower.
they all suck
HN's desire to pretend the data pipeline doesn't exist or isnt meaningful is silly.
Read the full thread on Hacker News →
Related stories
- Hacker News · 1 points · 9 days ago
- Hacker News · 1 points · 5 days ago
- Hacker News · 3 points · 3 days ago
- Hacker News · 7 points · 11 days ago
- Hacker News · 2 points · 1 day ago
- Hacker News · 1 points · about 15 hours ago