Comparing LLM-as-judge vs TypeSafe Jev for agent guardrails: same rules, same agent, measured on cost, latency, calibration and coverage. - deepansh-saxena/jev-guardrails

2 points•deepanshsaxena•7 days ago•1 comment•

1 comment

deepanshsaxena7 days ago
I replaced the guardrail layer in an agent with Jev. Same accuracy, 47× cheaper, 30× faster.

The agent: mock carrier customer support, 7 subagents, 43 tools. The guardrails: topical scope, desk boundaries, and 25 behavioural rules: is it evasive, over-promising, ungrounded, pressuring a customer who asked to cancel. Judgment calls. No regex settles any of them.

Same agent, same 25 rules, same thresholds, 51 labelled cases. The only thing I changed was what answers the question: a chat model reading each turn and reporting back, or Jev answering one yes/no per rule.

Accuracy came out even. That's the result that makes the rest matter. Everything below is what you get at the same quality.

One reply, 25 rules, both backends: 194ms vs 5,825ms. $125 vs $5,894 per million reviews.

The shape of it: deterministic rules belong in code. Auth levels, spend caps, a Luhn check don't need a model. But judgment calls are classification, and paying a generative model to classify buys higher latency and higher price.

Same accuracy, 47× cheaper, 30× faster.

Read the full thread on Hacker News →

Related stories