evaluation

12 stories and discussions about evaluation, aggregated from every source we track.

1.

We conducted cyber evaluations of Anthropic’s Claude Mythos Preview and found continued improvement in capture-the-flag (CTF) challenges and significant improvement on multi-step cyber-attack simulations.

18 points•martinald•6 months ago•1 comment
2.

Hello, I'm Shrijith Venkatramana, and I'm building LiveReview — a blast-radius aware AI code review...

11 points•shrsv•11 days ago•1 comment
3.

This informational PEP describes our shared view of the public C API. The document defines:

6 points•BiteCode•almost 3 years ago•0 comments
4.

An agent evaluation fails in CI but passes locally. Before blaming the model, ask whether both runs...

4 points•raju_dandigam•9 days ago•0 comments
5.

Context, limitations, and evaluation criteria for OliverDB performance results.

4 points•pnr-consulting•11 days ago•0 comments•
6.

<p>Includes results for core i9 7900x: <a href="https://github.com/intel-go/avx512counters/blob/master/avx512_core_i9_7900x.csv" rel="ugc">https://github.com/intel-go/avx512counters/blob/master/avx512_core_i9_7900x.csv</a></p>

4 points•quasilyte•about 8 years ago•0 comments
7.

We’re partnering with Accenture on independent evaluation of frontier AI—part of our recent commitment to embed evaluators at Anthropic. Both we and Accenture expect to invest at least $1 billion to build capacity in…

3 points•theanonymousone•11 days ago•1 comment•
8.

Workflow pattern using one LLM for generation and another for evaluation feedback loop.

2 points•softwaredoug•7 days ago•0 comments•
9.

LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly. We study JEV-as-a-Judge: evaluation with JEV, a decision-only judge that returns label probabilities instead of text, and whose confidence…

1 points•nico•about 23 hours ago•0 comments•
10.

Use GPT-6 Sol with AI SDK evaluation to review application data, interpret typed answers, and compare Astra and Luna with code examples.

1 points•flashbrew•4 days ago•0 comments•
11.

Introduction # This is a post about writing fast software in a roundabout way. Starting from a function definition, we define and synthesize a hardware circuit …

1 points•rgreen•4 days ago•1 comment•
12.

Learn how prompt evaluation noise can make a small prompt tweak look better than it is, even with temperature 0 and provider pinning.

1 points•khuyentran1401•9 days ago•0 comments•

Related topics