241 points•nicowaltz•1 day ago•94 comments•

94 comments

sharih1 day ago
What is the point of this, if it is p90 17 seconds? Might as well use an LLM. The beauty of Jev is that it is dirt cheap and insanely fast.
zihotki1 day ago
I would hold your horses to paint it as dirt cheap.. In my cases for spam detection Luna was 20% cheaper due to prompt caching, although not as fast.
nico1 day ago
For email you can use a classifier

One way: separately embed sender, recipients, subject, body - then use the embedding vectors as input to a logistic classifier

With that setup, I get 95% accuracy on email classification, training on 50-100 base examples. The model trains on CPU in under 1min, and it does inference in under 20ms (most of it is running the embeddings, so you can make it faster if you train your own embeddings model)

Here’s a gist with some sample code: https://gist.github.com/nicobrenner/056a5aaff5d0119c0032ecda...

That code applies the embeddings + classifier setup on the Banking77 dataset. It gets 93-94% accuracy depending on the embeddings you use (SOTA for this is ~95%, with much bigger and slower models)

atombender1 day ago
> hold your horses to paint it as dirt cheap

For a moment I thought this was going to be a metaphor — maybe an ancient Chinese proverb about how paint brushes are made from horsehair and how you can't hold the horse to paint before you've turned the hair into a brush.

calebhwin1 day ago
How are you benefiting from prompt caching for simple classification?
jedberg1 day ago
Are you getting better performance from an LLM than a Bayesian classifier?
HawtAds1 day ago
How many requests per second do you have for spam that you are reliably hitting the Luna cache?
amelius1 day ago
Next step: make it classify the next word.
esafak1 day ago
Jev ought to offer a flex mode that uses their spare capacity for a discount.
TN1ck1 day ago
I just did a run with a benchmark I just used to test other models against. (It's about detecting irony in german soccer tweets). On my M5 Pro with 48GB it took over 30min to decide on just 100 tweets, the thinking definitely takes long.

It performed quite below Jev, but above other open decision models I tested (68 correct vs 79 correct for Jev - see [1]). I'm running it for the moderation benchmark as well, but that will probably take a few hours on my machine.

[1] https://tn1ck.com/blog/jevdit

robbie-cabout 18 hours ago
One of Nico's colleagues here, I'm working on an MPS port https://github.com/PostHog/jeeves/pull/1 which should speed up M5 Pro performance quite a bit
TN1ck1 day ago
Update: Jeeves took about 2 hours to moderate 394 data points and performed really well. It’s not as good as Jev, but it’s super close! In general, it’s super cool that you can tune how strict you want content moderation to be with these models.
nicowaltz1 day ago
cool to see!
thm1 day ago
Ask Jeeves - Only took us 30 years to come full circle.
rsingel1 day ago
Too true. I worked there.

Ask Jeeves hired hundreds of cheap liberal arts majors to classify data, some users thought Jeeves was real, the stock spiked when big companies hired Jeeves to automate support thinking it was a silver bullet, and the whole thing collapsed when a better model came along, and it degenerated into ripping off rubes with bottom of the barrel ads.

kridsdale11 day ago
Sounds like the story of OpenAI in 6 years
NetOpWibby1 day ago
Damn, what a way to go.
victordmor1 day ago
I met one of the founders once in Oakland. Amazing fella.
kkukshtel1 day ago
You found the joke!
tmnstr851 day ago
this was the comment i came here for
onaclov20001 day ago
My bots are all named Jeeves lol. I have a CLI tool I use that connects up to a LLM I made and I call it Jeeves too ...so funny. I really didn't use Jeeves all that much I tended to use...I think it was called Web crawler pre-google era
davedigerati1 day ago
lol was thinking the exact same, named some ML projects Jeeves along the way...
itzikkatz1 day ago
Cool engineering, but 17s p90 latency kind of defeats the point of a Jev-class model, which is supposed to be fast and cheap. Losing 10 points on MMLU along the way doesn't help.
This reminds me of "on-premise cloud".

Read the full thread on Hacker News →

Related stories