Evaluates typed decisions (choice, score, noul) over 100+ languages in a single forward pass with calibrated probabilities. Outperforms TypeSafe Jev.
318 comments
OPs “marketing” is a single post on Reddit titled “ Predicting sales conversion probability from conversations using pure Reinforcement Learning”. Can you understand what that means? I can’t, and I consider myself reasonably technical. Is it obvious it has the same implications as Jev? Again, no idea. And it was just a single post on a subreddit that I don’t even browse! I see people on this thread saying “Jev is just BERT”. Sure, and Dropbox is just a ftp account mounted with curlftpfs!
I do feel bad for the author for finding something cool and being unable to brand it. But the full definition of “product” INCLUDES being able to coherently communicate it. In some sense the branding is just as much the “breakthrough” as the model.
Statistical modeling, from simple classical stuff up to modern deep learning, just has this dynamic where the theory is rich and bottomless, but the actual components of implementation are pretty neat and compact. So for any given idea, there are probably 20,000 other people who have had the same intuition, just with subtly different application or implementation. Add in that depending on what your particular flavor of research is, you might name an almost identical implementation something completely different. And it leads to a huge amount of sour grapes whenever anyone's idea really garners attention.
If you listen to any podcast with a founder in the ML space who has been in it for long enough, they will invariably say at some point "We actually developed xyz over a year before OpenAI"
I can understand it, and it wouldn't excite me at all.
Jev has a beautiful API and is advertised as something much more general.
(the project before it was rehashed into Laya since Jev was released)
It is arrogant and entitled for the author to take credit for the concept of RL over sequence embeddings, and none of the work that went into pretraining, not to mention the egregious target leakage [1]
[0]: Author fails to grasp the concept of virtual environments https://www.reddit.com/r/LocalLLaMA/comments/1kl0uvv/comment...
[1]: his `train.py` has `outcome` as a model input (conversation_metrics built from _parse_conversation which includes outcome): https://huggingface.co/DeepMostInnovations/sales-conversion-... https://huggingface.co/DeepMostInnovations/sales-conversion-...
[2]: 100% of this post is AI-generated https://www.pangram.com/history/97e0be84-391d-46b8-9c16-2d8f...
1)`outcome` is part of `metrics` at https://huggingface.co/DeepMostInnovations/sales-conversion-... and https://huggingface.co/DeepMostInnovations/sales-conversion-...
2) `metrics` goes into `ConversationState` at https://huggingface.co/DeepMostInnovations/sales-conversion-... and https://huggingface.co/DeepMostInnovations/sales-conversion-...
3) `metrics` (including `outcome`) makes its way into `ConversationState.state_vector` at https://huggingface.co/DeepMostInnovations/sales-conversion-..., and is returned from environment `step()` and `reset()` functions at https://huggingface.co/DeepMostInnovations/sales-conversion-... and https://huggingface.co/DeepMostInnovations/sales-conversion-...
4) model ingests `state_vector` as input at https://huggingface.co/DeepMostInnovations/sales-conversion-...
I was curious about this so I skimmed the paper [0]:
> SalesRLAgent achieved 96.7% accuracy, outperforming the best commercial alternative by 23.7 percentage points and the best LLM approach by 34.7 percentage points.
For a fuzzy natural language task like this, this magnitude of improvement should already set off alarm bells (Though i admit I'm not even sure what accuracy is even measured here, and the paper doesn't help either). Also, "best LLM" here refers to GPT-4 (at the time of upload, the public already had access to GPT-o3 and). I would have loved to contextualize the performance by looking at model size, but the paper is frustratingly devoid of detail in that regard:
> The core of SalesRLAgent is a reinforcement learning architecture consisting of: • A state encoder network that processes Azure OpenAI embeddings and features • A policy network that estimates conversion probability based on the current state • A value network that estimates the expected cumulative reward • A meta-learning module that assesses prediction confi dence
Also:
> Beyond technical metrics, we evaluated SalesRLAgent in real-world sales environments through A/B testing. [...] After 90 days across 217 representatives and 12,433 con versations, we observed: • 43.2% increase in conversion rate for the test group
This would be a pretty huge result but the fact that this is just shoved into a single paragrpah with no further discussion on methodology, baselines and setup makes me very suspicious.
Also, the way highly empirical fields like ML work is that it could very well be the case that typesafe had to do a _lot_ of work to improve this one, and in this field it ends up different enough that they feel they are doing something entirely novel[1]. I am not endorsing that 100%, but that happens a lot even between academics. In many cases it is valid.
[1] For example, this guys implementation seems to have atleast one serious issue, as {solution to OLS} points out in a sibling comment: https://news.ycombinator.com/item?id=49770027
“I personally found that this sequential approach captured sales dynamics much more effectively than traditional classification models.”
that first person phrase stuck out to me, especially given it had plural versions on either side, the author never edited for clarity or consistency
We haven’t seen any of these copy cats play doom or street fighter for instance; just categorize email.
I imagine once the author cools down and evaluates on a broad harness of tasks he may find that his new thing has a lot of engineering work ahead.
It reminds. Me of Devin. Took a while to debunk. Not saying Jev is a fraud , but the gap between structuring typed output and playing a game involving logical interpretation of frames made of pixels, screams unstructured interpretation they made and forgot to mention.
[1] https://arxiv.org/abs/2507.18546
Classical machine learning has been, for the most part, and just by the nature of science, behind academic terms and difficult to engage with as a product.
Jev did really well with coining up “System One” models and defining a standard application interface plus core primitives that landed in the current paradigm of software development.
I think it’s sort of like how Cursor reinvented autocomplete back then as a different UX and suddenly everyone was just using it because of how easy the bar was to understanding it.
Lastly, timing is everything. Just as Cursor had a first mover advantage, despite ML Ops being a thing for a while, they managed to encapsulate the concept behind a “System One” black box that fits the existing mental model for building software and shipping a data contract in the right point in time where the cost of tokens has been an important metric to watch.
No matter how much we pretend, that's how a lot of abstractions work. Things that touch the real world can change; there's a risk that the change could be as something as simple as a bugfix to changing the underlying implementation but preserving a higher level goal; you generally want a human in the loop to make sure the semantics work out and everybody's agreeing.
The relationship aspect of a business has a lot to do with how effective it is at continuing to justify its core value in an easy and relatable way; especially so when the decision makers that front the bill may not be as engaged with the underlying machinery behind the why it works how it does.
"Breakthrough", "our research went in another direction" , "Two years in stealth", "System One thinking model", "Jev can't hallucinate", "RLCD","We are doing very cool stuff, but we will have to hire you to tell you", - these are some of the things that they said on their website on the launch blog.
I had used versions of bert to achieve the same functionality years ago. But to me it seems like they were able to trick the VCs with "can't hallucinate" etc.
To the above author, kudos for sharing your work and making it open. Something like this shouldn't be closed in the first place when it has been available for so many years
I have a dozen different things at work that are currently using LLMs as classifiers for different questions. I don’t have the time, data, or resources to fine tune a model for each of them.
I haven’t had a chance to plug in Jev yet (waiting on approvals), but if it has the general intelligence claimed in the press release, then Laya is in no way comparable for my use case, and whatever TypeSafe has done is a substantial innovation over the Laya paper.
I agree with you. I think the OPs pushback is emblematic of a larger reaction I've seen that is, at the very least, misinformed.
There are a lot of approaches that use a self-attention backbone for classifier-style outputs. You have structured generation libraries like SGLang and Outlines, but those basically give you guided generation on an autoregressive model. You also have a bunch of models that are non-autoregressive that try something similar. Older NLP stuff applies here, and there's newer stuff using diffusion transformers for this purpose.
But I don't think the Jev author has ever said that he's the sole human, alone in a vast sea of misguided researchers, who is interested in schema-guided classification? I think he said he found a novel way to train a model for this task that has much higher general intelligence at much lower cost than other approaches. Which is an exciting result with lots of applications if it bears out.
I think some people are just reflexively skeptical of anything that gets a lot of hype. Maybe that's fair. Things that are wildly successful and high impact also tend to get a lot of hype though, so it seems like a poor filter.
it's very similar to jev's api and runs locally - if you like it, you can try jev for your actual usecases.
For me the cool bit is that it's all in-context learning or whatever so you can use it in any domain with zero setup.
Maybe bert and co. could do all the same things before, but the way in which you use them is quite different and that helps a lot.
https://huggingface.co/MoritzLaurer/deberta-v3-large-zerosho....
I don't understand why we lept to accusatory and personal, nor do I understand where this connects with the article, nor do I understand the assertions if I ignore either of those two things.
The article claims non-hallucination, it makes sense, then there's just someone sort of hand-waving at it's obviously false and people dumber than you were tricked. Not sure what trope to invoke here. Chesterton's fence?
And they’re acting like their probability isn’t as hallucinated as any other LLM guess.
Read the full thread on Hacker News →
Related stories
- Lobsters · 59 points · 11 days ago
- Hacker News · 2 points · 4 days ago
- Show HN: JevBench, a reproducible benchmark for typed decision modelsbenchmarkheaven.comHacker News · 149 points · 9 days ago
- The web being slow is a choice (and it's a stupid choice)albanbrooke.comHacker News · 1 points · 7 days ago
- Hacker News · 1 points · about 16 hours ago
- Hacker News · 1 points · 9 days ago