Solving the game with reasoning, not reinforcement learning. An interactive walkthrough of AI agents running a research loop to beat my game.

3 points•dustinlakin•10 days ago•1 comment•

1 comment

dustinlakin10 days ago
One of the most interesting parts was having this benchmark setup available over the last couple of months. It has made each new release something fun I can spend my night tokens on. The newest models really show the consistency I hadn't seen until Fable 5.1 and Astra 6.

Around to answer any questions or hear any feedback on the visualizations.

Read the full thread on Hacker News →

Related stories