We ran Jev, a calibrated judgment model, against our production cross-encoder on 12,927 labeled pairs. Precision 34.8% to 46.4%, recall 63.8% to 77.6%, same cost.
0 comments
No comments yet.
Related stories
- Hacker News · 5 points · 8 days ago
- Hacker News · 1 points · 9 days ago
- Hacker News · 1 points · 5 days ago
- Hacker News · 3 points · 3 days ago
- Hacker News · 7 points · 11 days ago
- I Put Jev Behind a TLA+ Spec and Ran 1,680 Chaos-Tested Pharmacy Decisions. Zero Wrong Verdicts.dev.toDEV Community · 3 points · 11 days ago