Compare large language models on more than 80 benchmarks, from GPQA Diamond and SWE-bench Verified to FrontierMath and Humanity's Last Exam. The table gathers every result published in Epoch AI's Benchmarking Hub, with…
#benchmarks#table#since#daily#updated daily#model benchmarks#benchmarks since#model benchmarks since
0 comments
No comments yet.
Read the full thread on Hacker News →
Related stories
- Hacker News · 1 points · 3 days ago
- Best LLM for every budget, updated dailybestmodelforyourbudget.terrydjony.comHacker News · 180 points · 8 days ago
- Daily Geography Adventuresmaptap.ggHacker News · 1 points · 10 days ago
- Hacker News · 4 points · 3 days ago
- Hacker News · 3 points · 11 days ago
- Hacker News · 1 points · 11 days ago