evaluation
12 stories and discussions about evaluation, aggregated from every source we track.
We conducted cyber evaluations of Anthropic’s Claude Mythos Preview and found continued improvement in capture-the-flag (CTF) challenges and significant improvement on multi-step cyber-attack simulations.
Hello, I'm Shrijith Venkatramana, and I'm building LiveReview — a blast-radius aware AI code review...
This informational PEP describes our shared view of the public C API. The document defines:
An agent evaluation fails in CI but passes locally. Before blaming the model, ask whether both runs...
Context, limitations, and evaluation criteria for OliverDB performance results.
<p>Includes results for core i9 7900x: <a href="https://github.com/intel-go/avx512counters/blob/master/avx512_core_i9_7900x.csv" rel="ugc">https://github.com/intel-go/avx512counters/blob/master/avx512_core_i9_7900x.csv</a></p>
We’re partnering with Accenture on independent evaluation of frontier AI—part of our recent commitment to embed evaluators at Anthropic. Both we and Accenture expect to invest at least $1 billion to build capacity in…
Workflow pattern using one LLM for generation and another for evaluation feedback loop.
LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly. We study JEV-as-a-Judge: evaluation with JEV, a decision-only judge that returns label probabilities instead of text, and whose confidence…
Use GPT-6 Sol with AI SDK evaluation to review application data, interpret typed answers, and compare Astra and Luna with code examples.
Introduction # This is a post about writing fast software in a roundabout way. Starting from a function definition, we define and synthesize a hardware circuit …
Learn how prompt evaluation noise can make a small prompt tweak look better than it is, even with temperature 0 and provider pinning.