scores

13 stories and discussions about scores, aggregated from every source we track.

1.
205 points•vinni2•4 days ago•406 comments•
2.
6 points•hn_acker•1 day ago•2 comments•
3.

This is a submission for the Kaggle Benchmarking Challenge. What I Benchmarked My first...

5 points•jaredchuvn•7 days ago•0 comments
4.

DeepSWE audit finds AI coding agents optimize for imagined graders, not users—a hidden reward hacking pattern.

3 points•guardiangod•2 days ago•1 comment•
5.

A mention count scores 'named and ignored' identically to 'the one recommendation'. We read 63 stored answers a second way to separate them: what it cost, what it found, and what we refused to let a model decide.

3 points•cuescout•4 days ago•0 comments•
6.

Anthropic's new Sonnet model scores 56, just 2 points behind Opus 5.5 (max), but at the highest Output Tokens per Task we have measured

2 points•spenvo•2 days ago•0 comments•
7.

Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index to bring SpaceXAI into the top 4 AI labs. Coding Agent Index performance has also improved, overtaking GPT-5.6 Sol

2 points•wertyk•8 days ago•0 comments•
8.

Play daily games, text in your scores, play with friends, and make your own.

1 points•dvdhutch•6 days ago•1 comment•
9.

Generate brandable startup names, check live domain & Instagram availability, then get an AI brandability score for how much competition it already faces.

1 points•coreinch•6 days ago•0 comments•
10.

How well can AI read automotive parts diagrams? Compare benchmark scores and costs.

1 points•thomas_ma•7 days ago•0 comments•
12.

Get your 1–8 PSL score, 16 facial measurements, canthal tilt angle, and personalized styling plan. Free, no signup needed.

1 points•amaz89•8 days ago•1 comment•
13.

Markets, fundamentals, earnings, options, Wall Street, news and crowd sentiment, run through quantitative and valuation models — one clear research view on every company we cover.

1 points•nacer222•9 days ago•0 comments•

Related topics