Ollaya downloads and serves open decision models on your own machine. Typed, calibrated answers in milliseconds, private and open source.

613 points•Ardakilic•5 days ago•145 comments•

145 comments

pradn5 days ago
I'm not sure what this means for AI startups if their innovations can be copied by OSS so quickly (what, like 2 weeks?). There's "consumer surplus" for everyone, to borrow an economic concept. But we do ideally want some of the surplus to flow to the innovator, too. I know there were precursors, but that's fine - it's hard to have a totally novel idea in such a popular field. I don't know what the end game is for TypeSafe - they'd need to demonstrate perpetually better results, or compete in another axis: UX, support, custom solutions, etc. So much of the time, someone proving a concept, or it simply getting enough publicity, is enough for a "Cambrian explosion" of follow-ups and copies. Famously, that was true for "Attention is All You Need", and the general idea of "next-token prediction" being so powerful.

We've stumbled into general differentiable models..

redox995 days ago
Because what they did is kinda trivial. Its basically like the Dropbox comment really[0], except here you don't need petabytes of storage and infinite VC pockets.

After chatgpt everything in AI mostly became LLMs and building wrappers around them. It's like people forgot how to do ML.

To those of us who actually trained models back in the day, its kind of cute to see people wowed by a classifier. Yes, this is 0 shot and doesn't need training (most people wanting this would've used structured output, this is cool because it's cheaper and faster). But anyone with basic ML knowledge could've built this in a few hours.

The question is mostly why wasn't this productized. And it's interesting indeed that it took this long to become a finished product.

[0] https://news.ycombinator.com/item?id=9224

djfdat4 days ago
Theorizing, but I'm guessing that before LLMs, people weren't using anything for situations like these. As people started using LLMs, people's use cases grew, LLMs got slower and more expensive, and people lost the wow-factor and are now worrying about price. Great timing to launch a product like this, where certain use cases can be distilled down into something faster and cheaper.

There's probably other areas where people are using LLMs where a more tailored ML solution might work better.

CactusOnFire4 days ago
I am looking forward to some hotshot implementing a jev-style model and then looking like a cost cutting, performance boosting wizard by replacing it with a logistic regression.
andai5 days ago
> The question is mostly why wasn't this productized.

Probably because doing it wrong (using an llm in place of a classifier) is more profitable? (For the people selling inference.)

eadwu5 days ago
The answer for why it wasn't productized might just be pretty straightforward.

LLMs still are better than Jev at the task, just across the board slower.

Anyone who had a reason to try this already tried it (ads/recommendations) - back in 2023/2024 during the first fine tuning wave and it was accurately determined that it was not worth the effort, the results were more bogus than just using CoT, so frankly parallelism meant nothing if bogus * parallel = bogus.

So thrown into the dumpster and nobody really cared to revisit because it was already tried.

Pretty much sometime between then and now it somehow became the state where the tradeoff makes sense now.

coolThingsFirst5 days ago
Can you explain how 0 shot classifiers work how come it isn’t trained on my data how does it know when issue is urgent lets say?
Fordec5 days ago
I think the idea of a "feature startup" is dead. What used to be a niche subscription business is now an individual Epic level of work. The smallest viable business becomes what two or three years ago was a mid tier enterprise. It is no longer "look at this tool I maintain", but "we take this specific approach using these hundreds of tools merged together to solve a problem in a specific way that nobody is going to compete with. Not because they can't compete if they wanted to, but that the competitions approach diverges in fifty different chosen ways that they are targeting a different market segment essentially."

I adhere to the idea that this is software's "Tower of Babel" moment where everyone just fundamentally ships things in completely diverging architectures, because creating a ground up architecture is no longer something that needs to be avoided for an economically viable business mode that in the past two decades would have otherwise incentivized people into industry standards. In a world where "taste" is the focus, single ingredients in the recipe aren't enough.

totetsu5 days ago
Are you saying laya copied from jev, and released in two weeks? If so I don’t thinks it’s quite as simple a story as that. https://xtxinversexty.com/layas-prior-art-claim-is-absurd/
cgio5 days ago
It a paradox when the article is claiming the prior art is absurd, but then goes on to analyse the one side and compare it to another for which most of the values (except scaling the concept) are unknown. And even for scaling, it uses the first, pre-laya instance to judge the limited schema, while overlooking that Laya is just doing this scaling. Important to note that prior art is not having built the exact same thing.
kensai5 days ago
There is definitely more to the story. There is a huge financial interest for each side to discredit the other. Fact of the matter is, we still don't know who will prevail. These are cutting edge tech stacks and they were just released.
janalsncm5 days ago
Presumably the training recipe and training dataset itself cannot be easily copied in a week or two. So if they want to shut down these competitor models they need to make it obvious how they are better than them.
mattstir4 days ago
> I'm not sure what this means for AI startups if their innovations can be copied by OSS so quickly

This particular "innovative" concept already has a rich, open research background. What Jev appears to have done is scale that up a bit and isolate good training data, which results in a great product but not really something impossible to imitate. The only major difference currently is that the open source decision models need to be fine-tuned as they're not trained off of the entire internet yet.

fooker5 days ago
For everyone dismissing Jev's innovation as being trivial, no it's not.

It is definitely not the MNIST classifier you had trained in 2019.

The difference is that you only train it once and the modern LLM machinery sort of takes care of that with large contexts.

It's great that Jev proved this is a viable product. I'd expect a great many research innovations coming from making this work better/faster/cheaper, and around interfacing modern agents with it.

hodgehog115 days ago
Just because it is zero-shot does not mean that Jev is an architectural innovation, especially in the year 2026.

Anyone fitting MNIST in 2019 was already outdated by several years at the very least. GPT-2 was 2019! We already had zero-shot classifiers then. In fact, the paper for GPT-3 was literally

"Large Language Models are Zero-Shot Reasoners"

These kinds of zero-shot classifiers were already developed and used in-house for many years. They just weren't commercialized as a separate product, because anyone who could use an LLM proper could build layers around it to fulfill any classification task like this.

fooker4 days ago
To get a classifier that worked in 2019, you had to train a model. Either from scratch or from a starting point.

GPT2 was absolutely unusable as a classifier. Using GPT3 as a classifier cost a few order of magnitude more than what this thing is priced at, and the context window was a few thousand tokens.

> These kinds of zero-shot classifiers were already developed and used in-house for many years

This is like Google's favorite coping mechanism for falling behind at AI. "We had everything inhouse for several years, we didn't release it for $reasons."

> anyone who could use an LLM proper could build layers around it to fulfill any classification task like this

You missed the part where it costs more than two orders of magnitude lower :)

I'm not claiming there are major architectural innovations, but that's not the point. Once you prove there's a market, there's a cambrian explosion of innovations.

syntaxing5 days ago
Hah you’re probably dating yourself. Keras came out in 2015 and that was one of the early examples with Theano backend. You could train MNIST since 2015 pretty straight forward. But comparing Jev to image classification is an unfaithful argument. Comparing it to ELmo or BERT is analogously better.
fooker5 days ago
You missed the point - you had to train BERT or anything similar to get useful results out of it.

Now all you need is to give it more context along with your query.

george_max5 days ago
Has anyone actually seen better or the same results with Laya compared to Jev? From my experience, Laya performs significantly worse. It's less confident and often makes wrong decisions with more complex queries.
jonmagic5 days ago
I've been following jevbench twice a day for the past week and that's been a lot of fun. Latest update:

Rank System Score Public / sealed accuracy Evidence

1 decider-4b v2 64.13 83.5% / 34.7% Evaluator-run, offline

2 Jev 1.13 63.29 86.6% / 36.7% Evaluator-run API

3 JevK5 v0.2 62.04 85.3% / 33.1% Evaluator-run

4 Cygnet 12B 61.76 87.9% / 33.8% Evaluator-run, offline

5 Hopper 59.43 82.3% / 34.1% Evaluator-run

28 Kev 4B 36.14 66.2% / 22.4% Evaluator-run

41 Laya 421M 30.25 58.4% / 30.8% Evaluator-run

https://benchmarkheaven.com/jev-models

rubymamis4 days ago
Did anyone else notice the huge gap between scores on private vs public for ALL Jev-like models compared to LLMs (such as GPT Luna)? Doesn't it mean those models aren't generalizing so not very useful on data they haven't seen?
Havoc5 days ago
Amazing - was looking for some benchmarks around this earlier
philipodonnell5 days ago
What the best way to see how a homegrown version compares?
scronkfinkle5 days ago
Yes. JEV generalizes better because they probably have an enormous corpus and trained on it for a long time. Laya's out of the box model is much weaker. However, in the age of LLM's it's incredibly easy and cheap to generate large datasets to fine tune laya for your task, and the training loop is pretty quick and cheap too.

It's so easy that I question why I would ever pay for JEV when eventually I'll have done enough random things that I will also have a large corpus and likely a general model as well.

mtkd5 days ago
Isn't the point of Jev that it generalises better?

It's a fast classifier you can use out-the-box, ~1.5bn tokens is about $40 (I've been hammering it)

It just works ... a whole bunch of low-level/low-importance workflow stuff that was getting farmed out to small/fast LLM models now has a competitive alternative ... and bits that hadn't even been considered to go into some external descision/classifier service can be tested/deployed at ~$0.00003/req

I don't get this wall of negativity on it, it's genuinely innovative/useful tech ... would expect HN to be more positive, regardless of whether it's the absolute best execution

_menelaus5 days ago
If you're so inclined it would be easy, fast and cheap to distill Jev for your task.
cobanov5 days ago
Developer here. You're right, Laya is a lot weaker than Jev, especially on harder queries. It's a small model, so it's fast, but that's the trade-off. The open models that get close to Jev are much bigger, and running those is what I'm working on next.
adinb5 days ago
It doesn’t to be a ton bigger, 16k and reliable 8k would be a godsend. (I run at 2k)
mikodin5 days ago
What are the models? I am super curious in these as well
lgas5 days ago
This is just anecdotal and I might be doing it wrong but I made jev and laya versions of a simple semantic grep tool (https://github.com/lgastako/jevplay) and played with them a bit, and at first it seemed like laya was comparable (eg on queries like "this is a mans name" or "this is a womans name" on names.txt) but the more I played with it, eg. "this is a vegetable" on foods.txt the further the gap widened in favor of jev. Then I started trying variations of the query eg simply "mans name" and for the most part laya just fell apart and didn't return anything useful for a lot of stuff. I was hoping to find that laya was competitive because it's much faster to have the model running locally but it's just not, yet.
cjonas5 days ago
I've been testing, for my use case laya didn't come close to decider was just as good. My experience with the 3 models matches the result here

https://benchmarkheaven.com/jev-models

alex7o5 days ago
Guys I have a real q, what is the difference between an instruct based re-ranker and laya/jev I just don't see it.

Edit: One is that jev/laya are tuned to have better probabilities, but a reranker can be fine tuned to do that as well. And jev/laya use RLCD?

Swizec5 days ago
> difference between an instruct based re-ranker and laya/jev I just don't see it

Main difference is that laya/jev/et-al give you a zero-shot classifier that requires no training. You can prompt engineer your way to a quick fairly reliable cheap enough decision engine that you can use to iterate quickly (by prompt engineering).

Right now a lot of people are doing this with LLMs and it's too slow and expensive.

Imo the right iterative approach to productionizing these systems is something like:

    1. Build it with an LLM. Iterate on the prompt
    2. Start building a real-world dataset
    3. When the prompt works, turn it into a clear rubric for Jev or similar
    4. Keep iterating until desired accuracy achieved
    5. Use the real-world evals you've built to train a custom classifier fine-tuned to your needs
You now have a system that has produced useful results in production from the very beginning and by the end it's a reliable super cheap classifier that can make thousands of decisions per second.
janalsncm5 days ago
I don’t think that’s it. I sincerely doubt most developers are doing side by side comparisons of calibration quality.

OpenAI has a section on their embeddings model api page for zero shot classification. Of course you can choose an open weights embedding too if you’d like.

https://developers.openai.com/cookbook/examples/zero-shot_cl...

I think Jev wins on marketing and convenience. Most SWEs don’t want to talk about embeddings, cosine similarity, or precision/recall tradeoffs. They want something which plausibly works and is easy to use.

kakugawa5 days ago
Jev's value becomes more apparent when the task is a moving target. eg an auto-mode classifier.
c10o5 days ago
Exactly. Just think about a discord or twitch moderation bot that screens messages in real-time. When the streamer starts playing a game, the whole context switch in an instant and comments like «kill them» will suddenly have a whole different meaning.
avereveard5 days ago
Calibrated probability across multi task with zero shot I guess. A reranker is single task and tuning it make it even more narrow. And I guess some piping to make multiclass efficient since you cannot mask logprob for independent questions in the same output space without throwing calibration away.
alex7o5 days ago
If you mean that I can compare the probilities between different tasks fair. This is a thing you can not do with a reranker. This sounds cool but not the amount of hype we got cool
solaire_oa5 days ago
I installed it, I tried the examples, it works.... But forgive my lack of imagination... what is this useful for?

Like, their example is of classification for a support interface.... `refund_requested`. Pretty convenient bool given the example is about a refund- what if 99% of submissions don't ask about a refund? Also, is that user not a `churn_risk`? What could possibly qualify as a churn risk if not a user asking for a refund?

https://ollaya.dev/library/laya The examples suffer the same problem of why I'd prefer to use a string column vs an enum. Changing an enum means you need to update the db, using a string you can do whatever.

I'm not trying to be negative, I genuinely want to know about some practical examples (that don't require tons of backwards maintenance).

devttyeu5 days ago
I have a lot of semi-practical examples of how you can use this model wrapped in unix-ish tools - https://github.com/aurorainfra/grev (readme links to docs of each tool with some more or less practical examples)

Really I think "smart grep" is a pretty good one ('look for an error looking vaguely like this'). Also I think sql-based shell history + decision model is quite good to make the last 'which one of those choices is best fit given users past few commands' etc.

spaniard892775 days ago
Isn't it better to use an LLM to train modernbert or xgboost et al?
solaire_oa5 days ago
Ok, those are pretty decent examples, and clears up the utility a bit: speed and tokens. Some of it's still a bit iffy (e.g. `cutv 'email address' 'phone number' < examples/users.csv`, csv is already in columns), but I can see using it for some niche queries. Neat tool.

I very much appreciate your to-the-point, non-vibed README as well, ty for that.

motoboi5 days ago
Is for when you want an AI to make a decision. If you have been using gpt or claude or open source models for that, than it’s a way cheaper alternative.

And if you have not been, it’s for when you have to extract the context from text. When you have numbers or fixed options, it’s just a matter of code.

So if you find yourself having to decide if a given user comment is a refund_request, that’s for that.

It’s not perfect, you still have to fine-tune (or calibrate) using examples you have (and keep those examples updated over time). But it’s way better than trying to parse text with regexes.

colordrops5 days ago
If you don't want to spend a lot and want low latency, e.g. home automation. "It's cold and dark in here, do something about it", it will then turn on the lights and heater nearly instantly.
maskedpirate5 days ago
sounds reasonable but doesn’t feel natural in a way that I can’t explain easily

Read the full thread on Hacker News →

Related stories