scores
13 stories and discussions about scores, aggregated from every source we track.
This is a submission for the Kaggle Benchmarking Challenge. What I Benchmarked My first...
DeepSWE audit finds AI coding agents optimize for imagined graders, not users—a hidden reward hacking pattern.
A mention count scores 'named and ignored' identically to 'the one recommendation'. We read 63 stored answers a second way to separate them: what it cost, what it found, and what we refused to let a model decide.
Anthropic's new Sonnet model scores 56, just 2 points behind Opus 5.5 (max), but at the highest Output Tokens per Task we have measured
Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index to bring SpaceXAI into the top 4 AI labs. Coding Agent Index performance has also improved, overtaking GPT-5.6 Sol
Play daily games, text in your scores, play with friends, and make your own.
Generate brandable startup names, check live domain & Instagram availability, then get an AI brandability score for how much competition it already faces.
How well can AI read automotive parts diagrams? Compare benchmark scores and costs.
Get your 1–8 PSL score, 16 facial measurements, canthal tilt angle, and personalized styling plan. Free, no signup needed.
Markets, fundamentals, earnings, options, Wall Street, news and crowd sentiment, run through quantitative and valuation models — one clear research view on every company we cover.