Compare large language models on more than 80 benchmarks, from GPQA Diamond and SWE-bench Verified to FrontierMath and Humanity's Last Exam. The table gathers every result published in Epoch AI's Benchmarking Hub, with…

2 points•Bobono•about 5 hours ago•0 comments•

0 comments

No comments yet.

Read the full thread on Hacker News →

Related stories