In the Stage 2 report we shipped 20 feature tickets on Fizzy and promised to explore benchmarking the agents all on max-effort. Now we have run it: every model on the board, same tickets, reasoning turned all the way…

2 points•doppp•9 days ago•0 comments•

0 comments

No comments yet.

Related stories