Evaluates typed decisions (choice, score, noul) over 100+ languages in a single forward pass with calibrated probabilities. Outperforms TypeSafe Jev.

1362 points•nandakishor_ml•12 days ago•318 comments•

318 comments

johnfn11 days ago
It’s a tale as old as time — people don’t understand that marketing and branding are just as important, if not more so, than the product. Jev is exceptionally-well branded. Anyone can look at the webpage and understand it, and the implications, instantly.

OPs “marketing” is a single post on Reddit titled “ Predicting sales conversion probability from conversations using pure Reinforcement Learning”. Can you understand what that means? I can’t, and I consider myself reasonably technical. Is it obvious it has the same implications as Jev? Again, no idea. And it was just a single post on a subreddit that I don’t even browse! I see people on this thread saying “Jev is just BERT”. Sure, and Dropbox is just a ftp account mounted with curlftpfs!

I do feel bad for the author for finding something cool and being unable to brand it. But the full definition of “product” INCLUDES being able to coherently communicate it. In some sense the branding is just as much the “breakthrough” as the model.

calebkaiser11 days ago
This is also a really common thing in ML specifically. We joke about getting Schmidthuber'd, which is when Jurgen Schmidthuber (sometimes correctly) announces that he or one of his colleagues actually proposed your thing 37 years ago in a Japanese linguists journal.

Statistical modeling, from simple classical stuff up to modern deep learning, just has this dynamic where the theory is rich and bottomless, but the actual components of implementation are pretty neat and compact. So for any given idea, there are probably 20,000 other people who have had the same intuition, just with subtly different application or implementation. Add in that depending on what your particular flavor of research is, you might name an almost identical implementation something completely different. And it leads to a huge amount of sour grapes whenever anyone's idea really garners attention.

If you listen to any podcast with a founder in the ML space who has been in it for long enough, they will invariably say at some point "We actually developed xyz over a year before OpenAI"

hiddencost11 days ago
Schmidhuber rarely ever executed the idea correctly, which makes his claims particularly obnoxious.
robrenaud11 days ago
> “Predicting sales conversion probability from conversations using pure Reinforcement Learning”. Can you understand what that means?

I can understand it, and it wouldn't excite me at all.

Jev has a beautiful API and is advertised as something much more general.

verdverm11 days ago
The title doesn't reflect the content of the paper or project, which uses things like RAG and an orchestrator, so more than "pure RL"

(the project before it was rehashed into Laya since Jev was released)

tinyhouse11 days ago
You're right but it's not the full picture. It's much easier to market when you have a name brand behind you. Not sure the author would've done much better even if he messaged it better. It's like the difference between someone random saying something smart on Twitter and no one gives a shit and Karapthy saying the same thing and everyone talks about it. I'm not saying it in a bad way - those with clout around them earned the people's trust by doing something right. But it's not easy to get there and there are many people doing great things that get very little publicity if at all. Not to mention in this case Jev came from a startup that raised a lot of money and can spend it on good marketing.
XTXinverseXTY11 days ago
OP's was leaky slop from day one [0][1], as is his article [2]

It is arrogant and entitled for the author to take credit for the concept of RL over sequence embeddings, and none of the work that went into pretraining, not to mention the egregious target leakage [1]

[0]: Author fails to grasp the concept of virtual environments https://www.reddit.com/r/LocalLLaMA/comments/1kl0uvv/comment...

[1]: his `train.py` has `outcome` as a model input (conversation_metrics built from _parse_conversation which includes outcome): https://huggingface.co/DeepMostInnovations/sales-conversion-... https://huggingface.co/DeepMostInnovations/sales-conversion-...

[2]: 100% of this post is AI-generated https://www.pangram.com/history/97e0be84-391d-46b8-9c16-2d8f...

OceanKing11 days ago
For anyone else verifying, the target leakage appears to be as follows:

1)`outcome` is part of `metrics` at https://huggingface.co/DeepMostInnovations/sales-conversion-... and https://huggingface.co/DeepMostInnovations/sales-conversion-...

2) `metrics` goes into `ConversationState` at https://huggingface.co/DeepMostInnovations/sales-conversion-... and https://huggingface.co/DeepMostInnovations/sales-conversion-...

3) `metrics` (including `outcome`) makes its way into `ConversationState.state_vector` at https://huggingface.co/DeepMostInnovations/sales-conversion-..., and is returned from environment `step()` and `reset()` functions at https://huggingface.co/DeepMostInnovations/sales-conversion-... and https://huggingface.co/DeepMostInnovations/sales-conversion-...

4) model ingests `state_vector` as input at https://huggingface.co/DeepMostInnovations/sales-conversion-...

fxwin10 days ago
> not to mention the egregious target leakage

I was curious about this so I skimmed the paper [0]:

> SalesRLAgent achieved 96.7% accuracy, outperforming the best commercial alternative by 23.7 percentage points and the best LLM approach by 34.7 percentage points.

For a fuzzy natural language task like this, this magnitude of improvement should already set off alarm bells (Though i admit I'm not even sure what accuracy is even measured here, and the paper doesn't help either). Also, "best LLM" here refers to GPT-4 (at the time of upload, the public already had access to GPT-o3 and). I would have loved to contextualize the performance by looking at model size, but the paper is frustratingly devoid of detail in that regard:

> The core of SalesRLAgent is a reinforcement learning architecture consisting of: • A state encoder network that processes Azure OpenAI embeddings and features • A policy network that estimates conversion probability based on the current state • A value network that estimates the expected cumulative reward • A meta-learning module that assesses prediction confi dence

Also:

> Beyond technical metrics, we evaluated SalesRLAgent in real-world sales environments through A/B testing. [...] After 90 days across 217 representatives and 12,433 con versations, we observed: • 43.2% increase in conversion rate for the test group

This would be a pretty huge result but the fact that this is just shoved into a single paragrpah with no further discussion on methodology, baselines and setup makes me very suspicious.

[0] https://arxiv.org/abs/2503.23303

porridgeraisin11 days ago
I think their problem is more not being cited by the team at typesafe, as in general academic politeness. On the one hand you have the charitable assumption that they developed it independently. On the other hand, my opinion is that it is naive to expect companies to do that even if they took inspo from it, especially when this is a core product theme, and not just some supporting infra. They will of course market it as their own. If they ever release a technical report, they might cite it there, but there is no way their landing page and announcement tweet cites it.

Also, the way highly empirical fields like ML work is that it could very well be the case that typesafe had to do a _lot_ of work to improve this one, and in this field it ends up different enough that they feel they are doing something entirely novel[1]. I am not endorsing that 100%, but that happens a lot even between academics. In many cases it is valid.

[1] For example, this guys implementation seems to have atleast one serious issue, as {solution to OLS} points out in a sibling comment: https://news.ycombinator.com/item?id=49770027

dcow11 days ago
I can understand why the author feels bitter but it still feels juvenile to me. Certainly both Jev and Laya are based on the research of countless prior papers and academics. Diogo decided to build a product out of the concept. The author didn't. Publishing research papers and model weights is probably part of the problem--it feels academic. If you look at the author's profile they focus on applying AI to healthcare. Not selling general AI type safety to AI pilled companies and devs. There's a big difference there. Whether that's good or bad you can argue all day. But for the author to expect otherwise is pretty weird. I do applaud them for not stewing too much on it and trying to do something about it, though.
operaopera11 days ago
I believe his qualms were with the "hype" in Jev's announcement: specifically calling this kind of model a breakthrough, without crediting previous art, and keeping everything closed source.
threecheese11 days ago
The hype is kinda nuts; I use X for ML/LLM stuff, and I just can't get away from Jev - even in my Following feed. Even days later 75% of posts are about "how I use typesafe for cooking breakfast!" or Jev clones.
m3kw911 days ago
if is closed source, how does he know there isn't some breakthrough he doesn't know?
prodigycorp11 days ago
And how is laya previous art? The project was vibecoded and posted yesterday.

https://github.com/NandhaKishorM/laya/commits/main/

https://huggingface.co/convaiinnovations/laya/commits/main

tomsyouruncle11 days ago
I’m not filled with confidence when the author’s first paper takes an RL approach but then doesn’t use it to change the action taken in the next turn. Seems like simple classification would achieve the same end. And this quote from the paper isn’t overly reassuring:

“I personally found that this sequential approach captured sales dynamics much more effectively than traditional classification models.”

https://arxiv.org/pdf/2503.23303

verdverm11 days ago
this was the period of arxiv history that led to the new vouching system

that first person phrase stuck out to me, especially given it had plural versions on either side, the author never edited for clarity or consistency

vessenes11 days ago
Agreed. Another difficulty here is there are not good benchmarks for this new architecture yet, so it’s easy to potshot and snipe, where jev seems to be pretty broadly intelligent/at least have had a lot of rl in different domains.

We haven’t seen any of these copy cats play doom or street fighter for instance; just categorize email.

I imagine once the author cools down and evaluates on a broad harness of tasks he may find that his new thing has a lot of engineering work ahead.

hirako200011 days ago
The doom demo would have to be reproduced to confirm what their model is capable of. Oh but it's all closed source, so who knows.

It reminds. Me of Devin. Took a while to debunk. Not saying Jev is a fraud , but the gap between structuring typed output and playing a game involving logical interpretation of frames made of pixels, screams unstructured interpretation they made and forgot to mention.

ktimespi11 days ago
"Juvenile" is a weird label to assign to someone whose work is re-presented by someone else and not attributed properly. People here had a very different take on the Navier-Stokes situation XD
kamranjon11 days ago
It is really interesting to see this claim, because i thought the current theory was that typesafe actually repackaged the work from GLiNER[1] - which does seem to be a closer match, and their original paper[2] predates yours by several years. Curious if you had heard of it before? It is also open source[3] and I think also has some good usage.

[1] https://arxiv.org/abs/2507.18546

[2] https://arxiv.org/abs/2311.08526

[3] https://github.com/fastino-ai/GLiNER2

hmokiguess11 days ago
I think the biggest lesson with Jev was the one of communication and understanding for the broader audience, sometimes a lot about innovating involves repeating yourself and translating your own thoughts to an intended audience.

Classical machine learning has been, for the most part, and just by the nature of science, behind academic terms and difficult to engage with as a product.

Jev did really well with coining up “System One” models and defining a standard application interface plus core primitives that landed in the current paradigm of software development.

I think it’s sort of like how Cursor reinvented autocomplete back then as a different UX and suddenly everyone was just using it because of how easy the bar was to understanding it.

Lastly, timing is everything. Just as Cursor had a first mover advantage, despite ML Ops being a thing for a while, they managed to encapsulate the concept behind a “System One” black box that fits the existing mental model for building software and shipping a data contract in the right point in time where the cost of tokens has been an important metric to watch.

sigbottle11 days ago
Furthermore, it's not about the current innovation right now - if you sell yourself on a broader mission, your core product can evolve and change with it, and you're more selling yourself as the guy who will make that abstract vision possible no matter what.

No matter how much we pretend, that's how a lot of abstractions work. Things that touch the real world can change; there's a risk that the change could be as something as simple as a bugfix to changing the underlying implementation but preserving a higher level goal; you generally want a human in the loop to make sure the semantics work out and everybody's agreeing.

hmokiguess11 days ago
Well said, once you put it out there in the world, then it's also about how will the customer react to it and having to own that relationship going forward.

The relationship aspect of a business has a lot to do with how effective it is at continuing to justify its core value in an easy and relatable way; especially so when the decision makers that front the bill may not be as engaged with the underlying machinery behind the why it works how it does.

prometheus199211 days ago
I think the main gripe that people had with Jev and Typesafe was the language used when they launched. To me personally it seemed like a parody/con/shady at first.

"Breakthrough", "our research went in another direction" , "Two years in stealth", "System One thinking model", "Jev can't hallucinate", "RLCD","We are doing very cool stuff, but we will have to hire you to tell you", - these are some of the things that they said on their website on the launch blog.

I had used versions of bert to achieve the same functionality years ago. But to me it seems like they were able to trick the VCs with "can't hallucinate" etc.

To the above author, kudos for sharing your work and making it open. Something like this shouldn't be closed in the first place when it has been available for so many years

yojo11 days ago
Is this equivalent though? The Laya article ends with “ Treat Laya as a fast foundation model to specialize, not as an omniscient zero-shot oracle.”

I have a dozen different things at work that are currently using LLMs as classifiers for different questions. I don’t have the time, data, or resources to fine tune a model for each of them.

I haven’t had a chance to plug in Jev yet (waiting on approvals), but if it has the general intelligence claimed in the press release, then Laya is in no way comparable for my use case, and whatever TypeSafe has done is a substantial innovation over the Laya paper.

calebkaiser11 days ago
Jev seems pretty cool! I just got access and have only gotten to do minimal experiments, but I love this general area of research and it fills a very real need.

I agree with you. I think the OPs pushback is emblematic of a larger reaction I've seen that is, at the very least, misinformed.

There are a lot of approaches that use a self-attention backbone for classifier-style outputs. You have structured generation libraries like SGLang and Outlines, but those basically give you guided generation on an autoregressive model. You also have a bunch of models that are non-autoregressive that try something similar. Older NLP stuff applies here, and there's newer stuff using diffusion transformers for this purpose.

But I don't think the Jev author has ever said that he's the sole human, alone in a vast sea of misguided researchers, who is interested in schema-guided classification? I think he said he found a novel way to train a model for this task that has much higher general intelligence at much lower cost than other approaches. Which is an exciting result with lots of applications if it bears out.

I think some people are just reflexively skeptical of anything that gets a lot of hype. Maybe that's fair. Things that are wildly successful and high impact also tend to get a lot of hype though, so it seems like a poor filter.

dwa359211 days ago
hey, you might wanna try this? - https://github.com/deepanwadhwa/OpenDecision

it's very similar to jev's api and runs locally - if you like it, you can try jev for your actual usecases.

vasco11 days ago
In a world of agents, doing a BERT run takes about 2 hours from having an empty folder. Just a thought you could consider. Once you've done the first you can do the rest of them before the end of the work day.
wild_egg11 days ago
Last time I did anything with a BERT, you had to train or fine-tune. Is that not still true?

For me the cool bit is that it's all in-context learning or whatever so you can use it in any domain with zero setup.

Maybe bert and co. could do all the same things before, but the way in which you use them is quite different and that helps a lot.

prometheus199211 days ago
It depends on your usecase but the models do show general capabilities. check this model out.

https://huggingface.co/MoritzLaurer/deberta-v3-large-zerosho....

evrydayhustling11 days ago
We used to use BERT-based embeddings + semantic distance for classification / decision problems in new domains. There was a lot of interest at the time in these kinds of pre-generative but portable models -- Meta's Prophet was another example that came up a lot.
dominotw11 days ago
you forgot the main one "from the guy who invented chatgpt"
refulgentis11 days ago
As long as we're in a thread about people "tricking", what you're claiming was written, or a synonym thereof, or kinda-sorta-the-same-thing, is not written anywhere.
refulgentis11 days ago
"But to me it seems like they were able to trick the VCs with "can't hallucinate" etc."

I don't understand why we lept to accusatory and personal, nor do I understand where this connects with the article, nor do I understand the assertions if I ignore either of those two things.

The article claims non-hallucination, it makes sense, then there's just someone sort of hand-waving at it's obviously false and people dumber than you were tricked. Not sure what trope to invoke here. Chesterton's fence?

seizethecheese11 days ago
I was confused by the “can’t hallucinate” thing, because it sounded like BS but people were taking it seriously. I purposefully asked a stupid question sort of like “this can’t hallucinate because it only has one output and there’s a schema?”. Was disappointed to learn the answer was yes.
MisterMunchkin11 days ago
Yeah it’s hilarious, it definitely can hallucinate. Just because it can only hallucinate “A” or “B” rather than a whole paragraph, doesn’t mean it is suddenly more accurate.

And they’re acting like their probability isn’t as hallucinated as any other LLM guess.

mrbonner11 days ago
But that hallucination is reproducible so you can adjust the prompt. Unlike an LLM in which everything is wildly not deterministic.
0x45711 days ago
Why would you think "can't hallucinate" means "can't pick wrong probability of an option" ?

Read the full thread on Hacker News →

Related stories