jev

349 stories and discussions about jev, aggregated from every source we track.

1.

Everyone is talking about Jev - here it is in 25 lines of Python.

674 points•bashbjorn•8 days ago•210 comments•
2.

Ollaya downloads and serves open decision models on your own machine. Typed, calibrated answers in milliseconds — private and open source.

599 points•Ardakilic•5 days ago•145 comments•
3.

Fine-tunes of Qwen3.5 and Gemma 4 for zero-shot classification - firelex/jeff

569 points•firelex•2 days ago•222 comments•
4.

tiny Jev-like family of decision models built on top of Qwen3.5 you can train and run on your own - jaredpalmer/kev

455 points•tosh•10 days ago•199 comments•
5.

TypeSafe's Jev is a genuine breakthrough – snap judgments with calibrated probabilities instead of generated text. My bet is OpenAI is already figuring out how to copy it, and then embed it inside its own models where…

322 points•JohnBerryman•8 days ago•223 comments•
6.

Watch the AI decision model Jev play Pokémon Red in its entirety, live.

281 points•pancomplex•5 days ago•123 comments•
7.

Jeeves – Reasoning improves Jev-like decision models - PostHog/jeeves

241 points•nicowaltz•1 day ago•94 comments•
8.

Left-pad strings with TypeSafe AI's Jev. For reasons. - f/jev-leftpad

227 points•fka•10 days ago•85 comments•
9.

A Visual Guide to RNNs, CNNs, Transformers, and Calibration, with Hands-On Experiments on Accuracy and Efficiency

182 points•Anon84•1 day ago•9 comments•
10.

Turns Jev into a chatbot. Contribute to kyle-pena-nlp/jevchat development by creating an account on GitHub.

159 points•kp1197•10 days ago•44 comments•
11.

52 Jev-class systems tested on 534 decisions. Jev 1.13.0 leads JevBench v1.3.0 with 74.4; compare open-source, self-hostable and hosted options.

149 points•florianstandhar•8 days ago•39 comments•
12.

I was intrigued by Jev and the self-hostable projects appearing around it, such as OpenJev and SemIf . Reading about them introduced me ...

148 points•allanrbo•5 days ago•44 comments•
13.
113 points•flxflx•4 days ago•49 comments•
14.

Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation. - trycua/cua

93 points•frabonacci•11 days ago•10 comments•
15.
88 points•nandakishor_ml•14 days ago•15 comments•
16.

Contribute to Amal-David/awesome-jev development by creating an account on GitHub.

70 points•frostbyte7•8 days ago•19 comments•
17.

September 2026. Every number here is from the benchmarks, and bash experiments/bench.sh --no-record reruns them without an API key.

65 points•tgluck•1 day ago•15 comments•
18.

RLCD is a calibrated, schema-conditioned extension of pairwise reward modeling: Bradley–Terry becomes Plackett–Luce, and the reward model becomes Jev’s typed decision interface.

47 points•tnspacetime•7 days ago•5 comments•
19.

Agent evals and guardrails in one request. Built on Jev, Kev and Laya. - openlayer-ai/jevals

47 points•gbayomi•10 days ago•6 comments•
20.

Review behavior, not just diffs. Jev prioritizes human attention; OpenAI explains the changes. Local CLI + agent skill + GitHub extension. - egma-ai/jev-code-reviewer

36 points•namanbhulawat•6 days ago•37 comments•
21.

Experiments in bulk inference with TypeSafe’s new model

33 points•ianmcook•1 day ago•5 comments•
22.

Jev is a useful zero-shot classifier, but its probabilities can't be calibrated for your data. Calibration depends on your data distribution, which Jev never sees, so treat its outputs as scores and recalibrate them on…

27 points•alexmolas•7 days ago•40 comments•
23.

Catch code issues before they catch you. Contribute to Alurith/jeff development by creating an account on GitHub.

27 points•imalessandro•12 days ago•5 comments•
24.
25 points•kunggaochicken•7 days ago•2 comments•
25.

How Jev's typed probabilistic decisions, LangGraph workflows, and Tenuo task-scoped warrants combine in a dependency upgrade agent.

22 points•niyikiza•7 days ago•4 comments•
26.

Ask a question about your document and find the relevant passages. Paste text or upload a PDF, Word document, text file, or Markdown.

19 points•irs•1 day ago•13 comments•
27.

kicking the tires on jev with 2048. GitHub Gist: instantly share code, notes, and snippets.

19 points•ndyg•11 days ago•2 comments
28.

Anyone who has moderated a live chat knows how fast things go wrong. Someone says something bad and...

15 points•shricodev•about 12 hours ago•4 comments
29.

TypeSafe released Jev on September 15th, 2026 and called it a System One model. A System One model is...

14 points•erikch•8 days ago•2 comments
30.
13 points•effisfor•3 days ago•19 comments•
32.

Your favorite corners of the internet. Posts, links, and conversations moderated by Jev.

11 points•TN1ck•3 days ago•1 comment•
33.

We run Jev on 100 real AI tool calls for the annotator pipeline. The errors exposed problems in both the models and our benchmark.

11 points•arseny_info•9 days ago•0 comments•
34.

We ran Kev 4B, an open reproduction of Jev, against Jev 1.13 on 362 questions written after both shipped: accuracy, calibration, token counts and speed.

10 points•felix089•5 days ago•0 comments•
35.

The Jev model from TypeSafe AI introduces a System One approach, delivering structured data rapidly...

10 points•truong_an_cornduck•6 days ago•1 comment
36.

An ML engineer's read on TypeSafe AI's Jev: what a non-autoregressive System One model changes for production classifiers, where it fits, and how I plan to test it.

10 points•pansuriyakartik•6 days ago•6 comments•
37.

Gemma 4 E2B, E4B, 12B and a 26B-A4B fp8 build read by their label probabilities with vLLM on one TPU v6e chip, checked against the same read on an NVIDIA L4 and against Jev 1.13.0's published results. What fits one chip, how to read labels when vLLM on TPU returns only the top 32 log-probabilities, speed, cost, and why no 31B loads today.

10 points•xbill•6 days ago•0 comments
38.

What the independent measurements of TypeSafe's Jev found in its first eight days: arXiv preprints, GitHub evaluations and blog benchmarks, each traced to its primary source. Accuracy, calibration, speed, cost, failure modes, the prior art, the open alternatives, and what is still unmeasured.

9 points•xbill•7 days ago•1 comment
39.

Schema-valid is not content-correct · Part 1/3 It starts with a storyboard a...

9 points•jimmyliao•11 days ago•1 comment
40.

Plain Gemma 4 26B read by its label probabilities against DiffusionGemma's one-step read, both as community 4-bit (AWQ) builds on one EC2 L4, on 1,200 labelled examples and on the 3,880-record public suite where Jev 1.13.0 has published results. Pre-registered, with accuracy, calibration, calibration after 0 to 150 labels, latency and cost.

8 points•xbill•7 days ago•0 comments
41.

Jev-shaped (TypeSafe System One) classification wrapper over OpenAI-like clients: probabilities and confidence instead of prose - zhulinchng/jevper

8 points•czl_my•8 days ago•3 comments•
42.

Can we now build “expert systems” that are actually useful and scale?

8 points•schmuhblaster•10 days ago•0 comments•
43.

We scored 2,449 security findings with eight models. Every model inflated severity on average. Explore the results, context failures, and what unnecessary investigations could cost your team.

7 points•brene•1 day ago•1 comment•
44.

Put anything on the tracks and find out whether Jev would pull the lever.

7 points•lanyard-textile•11 days ago•2 comments•
45.

jev-sec-audit is a drop-in CLI tool and GitHub Action that audits your package.json for threats. It...

6 points•dhanushnehru•7 days ago•0 comments
46.

Classify text in SQL with prompt_jev(), powered by TypeSafe's Jev. 100,000 rows labeled in 40 seconds for $0.50 at frontier-LLM accuracy. Live on paid plans. | Reading time: 5 min read

6 points•theanonymousone•8 days ago•0 comments•
47.
6 points•theanonymousone•9 days ago•0 comments•
48.

Open-source semantic decision engine for structured state and documents

6 points•dwa3592•9 days ago•0 comments•
49.

Answer choices, watch it predict you, and see the score improve as you play.

5 points•sdrth•4 days ago•9 comments•
50.

Contribute to RoyWiggins/jevlang development by creating an account on GitHub.

5 points•roywiggins•5 days ago•0 comments•
51.
52.

Get the AI to guess the word. Type clues without using the word itself. Five words. 30 seconds each.

5 points•rohanm93•6 days ago•1 comment•
53.

Send a state and questions to Laya, SemIf, Kev, NanoJev and Jevlike, the open-source Jev-style decision models, and read the calibrated probabilities that come back. Served on Beam's managed endpoints.

5 points•Mernit•7 days ago•0 comments•
54.

For the last few weeks my timeline has been nothing but Jev. TypeSafe AI shipped it, and within days...

5 points•programmerraja•7 days ago•0 comments
55.
5 points•htk•8 days ago•0 comments•
56.

We ran Jev, a calibrated judgment model, against our production cross-encoder on 12,927 labeled pairs. Precision 34.8% to 46.4%, recall 63.8% to 77.6%, same cost.

5 points•dennispi•8 days ago•0 comments•
57.

Last week TypeSafe AI unveiled Jev, their first example of a new category of model that they are calling “System One models” (I’m with Maggie Appleton, I think “decision models” …

5 points•benwerd•9 days ago•0 comments•
58.

The best stories from Hacker News, ranked by score over time.

5 points•jrhey•10 days ago•0 comments•
59.

Lint JavaScript and TypeScript against plain-English project conventions with Jev. - zdenham/jev-lint

5 points•zdenham•11 days ago•1 comment•
60.

Note: Pricing in this article is as of September 2026. TypeSafe lists Jev at $0.042 per million input tokens with output tokens free, and every cost figure below is worked out from that. Check the TypeSafe docs for…

4 points•dennisjoseph•2 days ago•0 comments•

Related topics