jev
349 stories and discussions about jev, aggregated from every source we track.
Everyone is talking about Jev - here it is in 25 lines of Python.
Ollaya downloads and serves open decision models on your own machine. Typed, calibrated answers in milliseconds — private and open source.
Fine-tunes of Qwen3.5 and Gemma 4 for zero-shot classification - firelex/jeff
tiny Jev-like family of decision models built on top of Qwen3.5 you can train and run on your own - jaredpalmer/kev
TypeSafe's Jev is a genuine breakthrough – snap judgments with calibrated probabilities instead of generated text. My bet is OpenAI is already figuring out how to copy it, and then embed it inside its own models where…
Watch the AI decision model Jev play Pokémon Red in its entirety, live.
Jeeves – Reasoning improves Jev-like decision models - PostHog/jeeves
Left-pad strings with TypeSafe AI's Jev. For reasons. - f/jev-leftpad
A Visual Guide to RNNs, CNNs, Transformers, and Calibration, with Hands-On Experiments on Accuracy and Efficiency
Turns Jev into a chatbot. Contribute to kyle-pena-nlp/jevchat development by creating an account on GitHub.
52 Jev-class systems tested on 534 decisions. Jev 1.13.0 leads JevBench v1.3.0 with 74.4; compare open-source, self-hostable and hosted options.
I was intrigued by Jev and the self-hostable projects appearing around it, such as OpenJev and SemIf . Reading about them introduced me ...
Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation. - trycua/cua
Contribute to Amal-David/awesome-jev development by creating an account on GitHub.
September 2026. Every number here is from the benchmarks, and bash experiments/bench.sh --no-record reruns them without an API key.
RLCD is a calibrated, schema-conditioned extension of pairwise reward modeling: Bradley–Terry becomes Plackett–Luce, and the reward model becomes Jev’s typed decision interface.
Agent evals and guardrails in one request. Built on Jev, Kev and Laya. - openlayer-ai/jevals
Review behavior, not just diffs. Jev prioritizes human attention; OpenAI explains the changes. Local CLI + agent skill + GitHub extension. - egma-ai/jev-code-reviewer
Experiments in bulk inference with TypeSafe’s new model
Jev is a useful zero-shot classifier, but its probabilities can't be calibrated for your data. Calibration depends on your data distribution, which Jev never sees, so treat its outputs as scores and recalibrate them on…
Catch code issues before they catch you. Contribute to Alurith/jeff development by creating an account on GitHub.
How Jev's typed probabilistic decisions, LangGraph workflows, and Tenuo task-scoped warrants combine in a dependency upgrade agent.
Ask a question about your document and find the relevant passages. Paste text or upload a PDF, Word document, text file, or Markdown.
kicking the tires on jev with 2048. GitHub Gist: instantly share code, notes, and snippets.
Anyone who has moderated a live chat knows how fast things go wrong. Someone says something bad and...
TypeSafe released Jev on September 15th, 2026 and called it a System One model. A System One model is...
Your favorite corners of the internet. Posts, links, and conversations moderated by Jev.
We run Jev on 100 real AI tool calls for the annotator pipeline. The errors exposed problems in both the models and our benchmark.
We ran Kev 4B, an open reproduction of Jev, against Jev 1.13 on 362 questions written after both shipped: accuracy, calibration, token counts and speed.
The Jev model from TypeSafe AI introduces a System One approach, delivering structured data rapidly...
An ML engineer's read on TypeSafe AI's Jev: what a non-autoregressive System One model changes for production classifiers, where it fits, and how I plan to test it.
Gemma 4 E2B, E4B, 12B and a 26B-A4B fp8 build read by their label probabilities with vLLM on one TPU v6e chip, checked against the same read on an NVIDIA L4 and against Jev 1.13.0's published results. What fits one chip, how to read labels when vLLM on TPU returns only the top 32 log-probabilities, speed, cost, and why no 31B loads today.
What the independent measurements of TypeSafe's Jev found in its first eight days: arXiv preprints, GitHub evaluations and blog benchmarks, each traced to its primary source. Accuracy, calibration, speed, cost, failure modes, the prior art, the open alternatives, and what is still unmeasured.
Schema-valid is not content-correct · Part 1/3 It starts with a storyboard a...
Plain Gemma 4 26B read by its label probabilities against DiffusionGemma's one-step read, both as community 4-bit (AWQ) builds on one EC2 L4, on 1,200 labelled examples and on the 3,880-record public suite where Jev 1.13.0 has published results. Pre-registered, with accuracy, calibration, calibration after 0 to 150 labels, latency and cost.
Jev-shaped (TypeSafe System One) classification wrapper over OpenAI-like clients: probabilities and confidence instead of prose - zhulinchng/jevper
Can we now build “expert systems” that are actually useful and scale?
We scored 2,449 security findings with eight models. Every model inflated severity on average. Explore the results, context failures, and what unnecessary investigations could cost your team.
Put anything on the tracks and find out whether Jev would pull the lever.
jev-sec-audit is a drop-in CLI tool and GitHub Action that audits your package.json for threats. It...
Classify text in SQL with prompt_jev(), powered by TypeSafe's Jev. 100,000 rows labeled in 40 seconds for $0.50 at frontier-LLM accuracy. Live on paid plans. | Reading time: 5 min read
Open-source semantic decision engine for structured state and documents
Answer choices, watch it predict you, and see the score improve as you play.
Contribute to RoyWiggins/jevlang development by creating an account on GitHub.
Get the AI to guess the word. Type clues without using the word itself. Five words. 30 seconds each.
Send a state and questions to Laya, SemIf, Kev, NanoJev and Jevlike, the open-source Jev-style decision models, and read the calibrated probabilities that come back. Served on Beam's managed endpoints.
For the last few weeks my timeline has been nothing but Jev. TypeSafe AI shipped it, and within days...
We ran Jev, a calibrated judgment model, against our production cross-encoder on 12,927 labeled pairs. Precision 34.8% to 46.4%, recall 63.8% to 77.6%, same cost.
Last week TypeSafe AI unveiled Jev, their first example of a new category of model that they are calling “System One models” (I’m with Maggie Appleton, I think “decision models” …
The best stories from Hacker News, ranked by score over time.
Lint JavaScript and TypeScript against plain-English project conventions with Jev. - zdenham/jev-lint
Note: Pricing in this article is as of September 2026. TypeSafe lists Jev at $0.042 per million input tokens with output tokens free, and every cost figure below is worked out from that. Check the TypeSafe docs for…