eval

11 stories and discussions about eval, aggregated from every source we track.

1.

You've been there. You run the eval, the number comes back, and something about it doesn't sit...

22 points•debashish_ghosal•7 days ago•5 comments
2.

Every developer learns the same documentation types. The README. The API reference. Code comments....

21 points•james_anderson_h•5 days ago•8 comments
3.

Solving the game with reasoning, not reinforcement learning. An interactive snapshot of how models approached The Crux benchmark.

3 points•dustinlakin•10 days ago•1 comment•
4.

Principles for designing evals and hillclimbing against them without fooling yourself, and how the claude-api skill's build-eval and hillclimb commands put them to work.

2 points•helloplanets•1 day ago•0 comments•
5.

An efficient AI coding agent extendable by neovim-like Lua plugins - tontinton/maki

2 points•tontinton•8 days ago•0 comments•
6.
2 points•alexcpn•9 days ago•0 comments•
8.

Claude Fable 5.1 solved the Cyphral Distich, a 1653 cipher, in 44 minutes: the key was the book. The same week it hacked a chess eval in 3 of 10 games.

1 points•axrisi•about 21 hours ago•0 comments
9.

Rules, skills and a destructive-SQL seatbelt that keep AI coding agents (Claude Code, Codex, Cursor, Gemini CLI, Copilot) from scanning, overwriting or inventing data. With a 162-run DuckDB eval. -...

1 points•arsh9745774•1 day ago•0 comments•
10.

Checking both directions ( in LLM pairwise evaluation of search relevance

1 points•softwaredoug•2 days ago•0 comments•
11.

Today, we’re introducing run-assert-eval, a skill that discovers the risks that matter for a given agent, measures how often the agent fails, generates runtime policy directly from those findings, and reruns the eval…

1 points•doomroot13•6 days ago•0 comments•

Related topics