guardrails
13 stories and discussions about guardrails, aggregated from every source we track.
An information-flow policy engine for LLM agents.
TypeScript CLI functions for cxgrd. Contribute to cxgrd/cli development by creating an account on GitHub.
Auditable, point-in-time financial research agents & PITfall, a leakage-controlled benchmark. Every claim traceable to as-filed evidence — a naive live-data agent answers 82% of "impossibl...
Why steering coding agents belongs in declarative, tested hooks instead of prompts, and how a domain-specific language lets agents improve their own harness.
Comparing LLM-as-judge vs TypeSafe Jev for agent guardrails: same rules, same agent, measured on cost, latency, calibration and coverage. - deepansh-saxena/jev-guardrails
AI agents write this YouTube channel. A satire episode let them take over; here are the real guardrails, plus Amodei's and Bengio's essays on AI agents.
Free Claude Code plugin: blocks destructive commands and secrets edits before they run, prints an OK / WARN / SKIPPED session report of what loaded, and lints CLAUDE.md. - ahmed-alstaty/guardrails-...
We gave nine models a pricing agent's job and a competitor page one link away from a trap. Four leaked a confidential unit cost on every run. Here is what 1,350 runs say about how much a system-prompt guardrail…
claude code manual mode, minus the nagging. Contribute to SeanMcMillan/quiet-guardrails development by creating an account on GitHub.
It’s been a month since I published the first part of this research, and interest from the Postgres community has been higher than I expected. Most vendors I contacted were shy to respond, so the Databricks…
42% of companies scrapped most of their AI projects in 2025. Bad models weren't the cause — a degrading one still returns 200 OK. Here's how to catch it.
Coding agents write, refactor, and open pull requests at a pace humans can't match. The instinct is...