Comparing LLM-as-judge vs TypeSafe Jev for agent guardrails: same rules, same agent, measured on cost, latency, calibration and coverage. - deepansh-saxena/jev-guardrails
1 comment
The agent: mock carrier customer support, 7 subagents, 43 tools. The guardrails: topical scope, desk boundaries, and 25 behavioural rules: is it evasive, over-promising, ungrounded, pressuring a customer who asked to cancel. Judgment calls. No regex settles any of them.
Same agent, same 25 rules, same thresholds, 51 labelled cases. The only thing I changed was what answers the question: a chat model reading each turn and reporting back, or Jev answering one yes/no per rule.
Accuracy came out even. That's the result that makes the rest matter. Everything below is what you get at the same quality.
One reply, 25 rules, both backends: 194ms vs 5,825ms. $125 vs $5,894 per million reviews.
The shape of it: deterministic rules belong in code. Auth levels, spend caps, a Luhn check don't need a model. But judgment calls are classification, and paying a generative model to classify buys higher latency and higher price.
Same accuracy, 47× cheaper, 30× faster.
Read the full thread on Hacker News →
Related stories
- Hacker News · 1 points · 9 days ago
- Hacker News · 1 points · 5 days ago
- Hacker News · 3 points · 3 days ago
- Hacker News · 7 points · 11 days ago
- Hacker News · 2 points · 10 days ago
- Hacker News · 2 points · 6 days ago