Run typed option logits and autoregressive JSON locally with open models and WebGPU.

722 points•ilreb•13 days ago•296 comments•

296 comments

mmastrac12 days ago
If you want to try a _legit_ Jev implementation that matches (at least in my evals), the vLLM patch to turn DiffusionGemma into Jev is available.

On my DGX Spark I get very similar latency numbers, and it matches my evals + or - a few points on each test (DG wins some, Jev wins some, both show low confidence when wrong).

I ran the same evals against a Qwen36 and it clearly lost to both of them, so you are leaving both knowledge and instinctual reasoning on the table with any smaller models, FWIW.

https://github.com/vllm-project/vllm/pull/57250

Vetch12 days ago
DebertaV3's architecture and noising should be even better as a basis because it had a couple inductive biases (cross encoder, disentangled attention and RTD corruptions) that enabled it to have unmatched weight performance ratio on such tasks.

My gut tells me that a better approach to a calibrated 0-shot classifier than shoehorning DiffusionGemma would be starting from another gemma, T5GemmaV2. Take its encoder and do continual training on an RTD objective and a large relational synthetic data mix. Then finetuning (multi-annotator data will help calibration) on as many proper NLI datasets as possible. That still lacks the DebertaV3 disentangled attention's inductive bias, however.

Jev also has its calibrated predictions component which is important. Temperature scaling is probably the easiest first pass. But there's lots of sensible options to improve on that.

ModernBERT might be the easier, more stable starting point than T5Gemma though.

mmastrac12 days ago
I suspect you could get interesting results, but DiffusionGemma has a lot of knowledge that may be challenging to train into the smaller models. The advantage of pulling a fully-trained diffusion model off the shelf is that it already knows all of this, has been trained as a MoE, etc.

What I think these models actually need is a structured decision thinking mode. As it stands now, the only way to think about the answers with DiffusionGemma is to diffuse a thinking block, but giving the model an auto-regressive thinking space to reason, even just lightly, could drastically improve performance.

mungoman212 days ago
This is very interesting! Seems like a promising direction.

I wonder though if it supports the same claims as Jev: answers are not impacted by other answers to the same questions, nor the existance of other questions? It seems by sharing KV cache all questions will be visible. And I think the diffusion causes the answers to attend to eachother?

Also I think without fine-tuning we can not say that the probabilities it output are actually probabilities. Maybe fine tuning using Brier scoring would do the trick?

Maybe some kind of hierarchical structure of the KV cache can make questions independent, and with smaller diffusion canvas’ generated in batch can be a way to make answers generate independently?

cmrdporcupine12 days ago
> It seems by sharing KV cache all questions will be visible

Yeah, this is partially why in my approach I've done this instead, and not used diffusion model:

1. Convert the state into one shared prompt.

2. Run that shared prompt through the model once.

3. Fork the model’s internal state once per question.

4. Add a different question to each fork.

5. Ask each fork for its next-token scores.

6. Calculate only 64 possible label scores—not the whole vocabulary.

Basically ... skip decode.

Won't be as fast as doing diffusion model parallel across a pile of questions at once, but:

a) let's you use pretty much any existing text model (with some modifications). I've got qwen3.6 moe and qwen3.8 flash next running, am getting gemma4 working now

b) the problem you identified

It's possible I'm getting high on my own supply and misunderstand entirely the whole thing, but it seems to work?

https://github.com/rdaum/eider/commit/9c2d5c049068c33da2a48b...

I don't have the chutzpah to go creating PRs for vLLM to do the same.

mmastrac12 days ago
I believe you get better answers by diffusing each question together, but the PR's server gives you finer control over that. If each question is independent, you can get better parallelism.
cmrdporcupine12 days ago
This PR is interesting but it's making the assumption that what Jev has done is based on a diffusion model or that a diffusion model is superior for this work. Which may or may not be the case.

If I understand it though it does mean you can evaluate a bunch of questions simultaneously, which is an advantage.

Also: While I think it's expected/normal to see LLM-generated programs... there's a lot of LLM written comments in that PR, which is sad to see. Auto-human.

wuhhh13 days ago
I don't understand how this is different from oai "structured output" (and whatever the similar paradigm was on Sonnet ~3.7 back then) which everyone moved on from. On their gh they say:

"Jev is TypeSafe's closed service for runtime-defined semantic decisions. This project reproduces that interface pattern with open models; it does not reproduce Jev's undisclosed model or training"

As someone else pointed out it isn't actually Jev... can someone enlighten me

Topfi13 days ago
Jev is, as far as I understand, essentially very optimised for zero shot classification [0]. Something like BERT could be and has been tuned to provide similar "decision making" at a similar latency and cost advantage quite some time back. Advantage over full on LLMs is mainly the efficiency and of something like Jev over e.g. the encoder/decoder based classifier I had in front of an LLM to route to different prompts depending on the users likely needs, that Jev does perform at a more consistent level, allegedly roughly akin to GPT-5.6 Terra, but at the lower cost and latency. Currently testing that, but seems promising, if Jev classifies at or above Terra level, I see no reason not to leverage it.

Can add that I tried using a heavily pruned mt0 based model for structured classification along with structured output for local tagging and simple renaming suggestions. While it does work, the balance is hard to get right for the machine I was targeting as a minimum spec (Macbook Neo), so that's on ice. Focusing on one of the tasks easily goes below 100mb with solid latency across all EU Latin script languages, but the second you add a few, it's simply not in the quality budget, so while LLMs can do anything Jev and similarly focused models can, it comes at a literal cost. Could maybe accomplish the goal with multiple models (BERT+mt0+...), but that get messy.

In general just happy to see a bit of the millions flooding into the industry being used to improve on less flashy but immensely useful solutions. It's amazing that you can technically use LLMs for most tasks, but not every org has a near infinite budget and there is still a lot to gain from applying more recent learnings to old solutions along with just updating their training data to the current year. Also makes business sense, competition on frontier or mid-tier LLMs is vicious, focusing on an underserved niche with clear application is clever.

[0] https://huggingface.co/tasks/zero-shot-classification

orbital-decay13 days ago
It's a non-instruction-tuned classifier model trained on a confidence-aware RL variety that generates its own schema and follows it, with a confidence score output. Think BERT on crack, smart enough to be used as a decision maker (conceptually). They call it "not an LLM" because it's non-generative but of course it's a language model in the same way all non-instruction-tuned classifiers are.
mtkd13 days ago
I was a bit skeptical when read the initial pr on it, yesterday ran a test involving ~250M tokens, something we measure went from ~60% to >80% success (with almost no tuning) and at less than 50% cost the low-end LLM was running at, looking at it more seriously now ... the servers are US-only currently I understand and ZDR is by request
ozgung13 days ago
Isn't that the same transformer at the end of the day? It must be faster only because it generates a single token output, just one evaluation of the model. It takes the same input context and has the same O(n^2) attention blocks. It probably takes options as appended to the input and returns a probability over them instead of the whole dictionary. It's post-trained to do that specific job. If so what's the big deal?
3abiton12 days ago
This is a great evolution in the right direction compared to LLM + pydantic and temperature 0.
petesergeant12 days ago
No idea at all what this particular project called openjev is doing, but https://github.com/TheoLeeCJ/openjev and https://github.com/ekzhang/openjev-sglang (neither of which I have any relation to) generate a single token, rather than JSON structured output. I wrote up this technique here: https://sgnt.ai/p/jev/
wuhhh12 days ago
I came to say thanks for the link to sgnt.ai on Jev, then realised you're the author! Well, thank you so much, I feel informed :)
messh12 days ago
In Jev you pass options in the input and its output just gives some probability for each. Oai structured output just follows a schema. The exact output is still generated and there is no probability
brokensegue12 days ago
you can ask structured output for probabilities...not that they necessarily mean anything.
mritchie71213 days ago
in short: it's faster, cheaper, smart structured output.

each "question" is answered in parallel instead of a sequential (like an LLM). so if you have an input like:

    {"is_it_hotdog": noul, "is_it_apple", noul}

it answers is_it_hotdog and is_it_apple in parallel and gives a probability.
satvikpendem13 days ago
Can't I just parallelize my LLM calls myself for each question?
prodigycorp13 days ago
These one shot vibecoded sites are always a complete visual headache. Endless clutter, pointless filler text all over the place, and zero regard for actual usability.
dkarl13 days ago
AI output right now is like a final exam essay response from an anxious student. Instead of being edited for focus and clarity, it's anti-edited to cram in as many details as possible. Instead of worrying that the reader might get bored or confused, it assumes that the reader has no choice but to read the whole thing, even if they get a headache. It doesn't care about picking the most useful perspective on a problem; it cares about covering every possible angle that a grader might use to dock points from it.

It's basically the work you get from a smart, diligent person who is oblivious to any shared goal and approaches every assignment with a CYA attitude.

prodigycorp13 days ago
Pretty good analogy. I'd also compare it to a junior employee who tries to make people care about the how of their work rather than the results.
akoboldfrying12 days ago
That, and Every Sentence Is Punchy.

Every sentence sounds like it's trying to be in the trailer for a film.

chrismarlow912 days ago
Here's my prompt when I don't know what the hell the AI is saying.

"Expand. Clarify for human. 5 minute read max. Senior engineer audience."

ikari_pl12 days ago
You just very nicely explained why I sometimes overcommunicate.... That's exactly how it feels
TeMPOraL12 days ago
It's not a bad attitude given that I don't trust it did the work right, so the more details it can include, the more opportunities I have to spot errors, which invalidate the whole result - and conversely, if all details seem right and self-consistent, it gives me greater confidence the work is correct.
monkeydust13 days ago
Berkshire got it right a long time ago.

https://www.berkshirehathaway.com/

phoghed13 days ago
Nah, if this was the OP website you’d be complaining that it tells you nothing and you have no idea what they do or what they are presenting still. Also that it looks like shit on mobile. You’re just glazing the company in this case.

You should go with the canonical HN quality website references: McMaster-Carr, Craigslist

Topfi13 days ago
> If you have any comments about our WEB page, you can write us at the address shown above. However, due to the limited number of personnel in our corporate office, we are unable to provide a direct response.

A profoundly polite way to tell someone to stuff it.

m12k13 days ago
This site proves to me that the better you are at the things that matter most in your niche, the more you can get away with not even trying in other areas.
BrokenBuild13 days ago
this is a great example for me to use in meetings. I often see people looking for "good" examples of web design from fortune 500 companies or similar. Gonna use this to throw a wrench in that one soon.
justinhj12 days ago
The Geico ad gave me a laugh.
binlog12 days ago
There's a "unsloppify site" toggle on top but the unsloppified version looks exactly as vibe coded as the regular one.
jamilton12 days ago
Yeah, I'm not sure which way is supposed to be "sloppified". The default looks stylistically less slop-like, but obviously has the same filler content issue.
ljm12 days ago
I love how the 'unsloppify' button changes the theme but nothing else, and also messes with the layout enough that you can't actually toggle it without scrolling and repositioning your cursor.
assimpleaspossi13 days ago
I had to re-read a few times to figure out what the site was for and about. Still not sure I understand but that's the problem for them. If I'm a customer, I'm gone cause I can't figure out what it's for and I see this far too often nowadays for a lot of technical sites.
Hackbraten13 days ago
I already closed it after two pageful of not explaining what this thing is.
kul_13 days ago
Is it only me or do others also find LLM generated websites so off-putting?
djaro13 days ago
Unsolvable problem.

Why was the aesthetic standard to be pale when workers worked the fields and royals were inside, but tan when workers moved into factories and only the rich could afford to go on a beach vacation?

Aesthetic standards are formed by association. Its why sites that are "well designed" but obviously just use a squarespace or wix template feel so cheap. Why millenial flannel went from hip to standard to outdated. Why purple was the color of royalty before we could synthesize the pigment.

Having good design is about associations. Whatever design LLMs will default to, it will always feel cheap because we will learn over time that that design means cheap. Having good taste is about being ahead of the curve. An LLM cant be ahead of the curve because then that becomes the standard, and theres a new ahead.

You can use LLMs to make novel looking websites by carefully telling it to add certain details, use certain elementd, etc. At that point youve looped back to being a graphic designer.

ncphillips13 days ago
I think what you’re talking about is real, but it’s only part of the problem. The issue is it’s poor design. There’s a lack of consistency that is really off putting. Spacing is inconsistent and doesn’t create a sense of visual hierarchy. Buttons, inputs, selects, call-outs, table cells are barely distinguishable from each other, but also inconsistent within their own categories. The copy is also confusing. I don’t even know what this does.
miki12321113 days ago
This is, once again, about diversity and the lack thereof (and I don't mean diversity in a political sense).

LLMs seem fundamentally incapable of producing truly diverse outputs, truly creative and different responses to the same prompts in different runs. Because you and me use the same Claude, if you want a website and I want a website, we'll get (almost) the same website. This is not some BS about "the average of its training data", most of the LLM style (both in design and in text) comes from reinforcement learning. You could RL Claude to produce a very different style, but you couldn't RL it to produce a different style for me than it does for you.

I think this is also where a lot of the complaints about "Claude writing" come from.

zaep13 days ago
I don't know, I think you have a point that aesthetic preferences are subjective and shifting. But there is also all-caps monospaced text with emdashes in it on the site; just an example of something I think would not turn into a fashion at any point because it just looks silly (subjectively, to me at least). Thus I don't think the antipathy of people towards these llm-generated landing pages is entirely based on associating it with other LLM sites, there is at least an element of it clearly not being through as much human review and interaction as a hand-crafted landing page necessarily would be.
__rito__13 days ago
> Its why sites that are "well designed" but obviously just use a squarespace or wix template feel so cheap.

Same. I honestly am very satisfied with the aesthetics of free Wordpress blogs. Like Terry Tao has. I also have one.

sim04ful13 days ago
This is a problem that i'm actively working on (https://fudge.design), what i've realised is that it's simply not an issue of capability - given a well crafted site and a competently written visually aware harness, most recent models can replicate that website.

So it's what lies between saying "I want x website" -[.....] -> Code+Assets

The issue has to do with specification fidelity, in short a grill-me style aesthetic interrogation using illustrative tooling - ascii diagrams for specifying layout, copy and user-flow, image-gen mockups for higher fidelity mockups. References are also very important for nailing down the aesthetical qualities. I've noticed it's far better vs purely text description to simply gather up a mood-board telling the llm to find commonalities and come up with a design system and brand guide.

So I don't believe it's an unsolvable problem, it's simply a lack of effort on the implementors part. Also there's probably some survivor's bias here (you won't notice an intentionally designed vibe-coded site)

For example here's one reference exploration site i recently made with grok: https://explorer.withfudge.com/

joegibbs13 days ago
The problem is the overabundance of text, they can’t let it breathe. Everywhere has to be filled up with bits of hardly-relevant text.
sheepscreek13 days ago
Language models, amiright?! Text is the blood flowing through their veins. It’s all they care about.

As they say, to a hammer, everything is a nail.

testycool13 days ago
I tell the agent to outline the greebling, which is the when you add extra details that aren't really necessary.

Initially it feels like the result will be too empty, but once the greebling is removed it most often looks better

berofeev12 days ago
I find the models conceptually design the structure, then fill it up with text.

This leads to 'fixing' the amount of text areas it needs to fill, so it works to a constraint of having to fill a collection of text areas instead of outputting the message that would otherwise best suit.

My way to combat this is to start by refining the words/messages to be put on the canvas before letting the agent try designing something.

yellowapple12 days ago
Honestly I prefer that over the overabundance of empty space that's been the norm in “modern” web design for more than a decade now.
phoghed13 days ago
I appreciate a nice brutalist aesthetic like this tbh. It’s also good that there’s a baseline for quality in terms of layout and spacing and contrast and whatnot usually, so the HN webshit meta conversation has shifted from that to whinging about an LLM making it.

The overall arrangement and useless shit LLMs put in the copy is often annoying though.

nz13 days ago
This site actually reminds of the TUIs that one uses to install an OS from the text-console. It's not so bad. The prose itself is irritating. The site itself also has some bugs (text overlapping with UI borders for no reason). The lime-green color is a little awkward to my eye, but maybe that's just me (I say this as someone who usually likes lime-green -- maybe the problem is that this site needs _more_ lime-green).
alex_suzuki13 days ago
Same. I can’t really put my finger on what exactly is turning me off though. I mean, apart from the obvious AI-generated text.
adventured13 days ago
It's overly automated and repetitive in its styling. Humans make odd stray adjustments to styling manually. LLMs build pages very efficiently. Unless you're very anal-retentive when building a site, there's going to be some distinct flair that isn't just a repeating segment.

It's like it was made by the world's most anal-retentive Wordpress theme builder. They went over it a thousand times until it was perfectly optimized, no distinguishing marks, no stray tiny misalignments, no single-use stylings.

sheepscreek13 days ago
Mainly cause you never know what you’re getting. Over time, we trained our minds to believe that a well put site = effort, so at the very least people behind it cared. Now, it takes zero effort to make a site look good. So appearance in general means even less. In fact, now a poorly put together site might mean someone cared, wrote it by hand, flaws and all, to give you the human to human experience.

If there is a silver lining in all this, this might get us to appreciate the flaws in all humans, heck even yearn for them.

Read the full thread on Hacker News →

Related stories