calibrated
7 stories and discussions about calibrated, aggregated from every source we track.
Jev is a useful zero-shot classifier, but its probabilities can't be calibrated for your data. Calibration depends on your data distribution, which Jev never sees, so treat its outputs as scores and recalibrate them on…
Jev promises decisions with honest probabilities in milliseconds. How does it work, and how close does a stock open model get?
Evaluating Jev's calibration on two hard datasets.
Build calibrated AI classifiers from human feedback using Jev and GEPA. - sutro-sh/jev-align
We ran Jev, a calibrated judgment model, against our production cross-encoder on 12,927 labeled pairs. Precision 34.8% to 46.4%, recall 63.8% to 77.6%, same cost.
Under 120 MB. One forward pass returns typed, calibrated answers. By toolchain.studio.
Detectors of alignment failures screen deployed language models and score alignment benchmarks. Most are generative judges that spend a decoding pass on every criterion, and classifiers that read token probabilities,…