eval
11 stories and discussions about eval, aggregated from every source we track.
You've been there. You run the eval, the number comes back, and something about it doesn't sit...
Every developer learns the same documentation types. The README. The API reference. Code comments....
Solving the game with reasoning, not reinforcement learning. An interactive snapshot of how models approached The Crux benchmark.
Principles for designing evals and hillclimbing against them without fooling yourself, and how the claude-api skill's build-eval and hillclimb commands put them to work.
An efficient AI coding agent extendable by neovim-like Lua plugins - tontinton/maki
Claude Fable 5.1 solved the Cyphral Distich, a 1653 cipher, in 44 minutes: the key was the book. The same week it hacked a chess eval in 3 of 10 games.
Rules, skills and a destructive-SQL seatbelt that keep AI coding agents (Claude Code, Codex, Cursor, Gemini CLI, Copilot) from scanning, overwriting or inventing data. With a 162-run DuckDB eval. -...
Checking both directions ( in LLM pairwise evaluation of search relevance
Today, we’re introducing run-assert-eval, a skill that discovers the risks that matter for a given agent, measures how often the agent fails, generates runtime policy directly from those findings, and reruns the eval…