In the Stage 2 report we shipped 20 feature tickets on Fizzy and promised to explore benchmarking the agents all on max-effort. Now we have run it: every model on the board, same tickets, reasoning turned all the way…
0 comments
No comments yet.
Related stories
- Show HN: Groundtrack – Continual learning for coding agentsgroundtrack.devHacker News · 1 points · about 15 hours ago
- Hacker News · 131 points · about 11 hours ago
- Hacker News · 45 points · 9 days ago
- Hacker News · 1 points · 3 days ago
- Hacker News · 37 points · 13 days ago
- Hacker News · 23 points · 3 days ago