Did you ever click on an “AI Arena” expecting glorious battle and instead get a boring benchmark? If so, this project is for you: proper life-or-death fights between four models on a picturesque 8×8 grid. May the most…

120 points•hp6•3 days ago•45 comments•
Did you ever click on an “AI Arena” expecting glorious battle and instead get a boring benchmark? If so, this project is for you: proper life-or-death fights between four models on a picturesque 8×8 grid. May the most intelligent one win!

Click on any of the matches to spectate them.

Code: https://github.com/hp6/ai-arena

45 comments

CuriouslyC3 days ago
This is similar to how I evolved the AI for my game Hexborne (Think X-Com meets Magic: The Gathering). I had agents designing sets of behavioral heuristics for bots (as "genes"), then enter them into 10k+ match tournaments in an iterative process. After each tournament agents could inspect all the heuristics and try to craft updated heuristics to improve their performance. Winning heuristics got a full statistical validation before being rolled into the baseline set that all agents build on top of.

To multitask I did all this over a multiplayer game server to harden the netcode and ferret out softlocks.

deadbabe3 days ago
It’s amazing how people with AI are discovering classic techniques for balancing games (i.e. Monte Carlo methods, genetic algorithms…), but somehow implementing them way less efficiently, and without mathematical rigor, basically just having an LLM do the work of deterministic math functions.
CuriouslyC3 days ago
You're being presumptuous, I explicitly was explicitly thinking of a GA when I set this up, and how is it inefficient when the AI reduces the number of non-viable policies (and thus the number of wasted simulations) by multiple orders of magnitude compared to randomized policy generation?
idjdiejciejfj3 days ago
…what makes you think people are implementing them less efficiently? I’d be inclined to agree with you but your reply reeks of “no true Scotsman” strawman arguments that it’s hard to take it seriously.
nizarmah3 days ago
There's a lot of information in this world. It's fun stumbling on that information sometimes. For me personally, it makes learning much easier and gradual, than reading about something from a book or LLM answers.
fragmede3 days ago
Someone hasn't seen xkcd 1053 recently enough
PaulStatezny3 days ago
The rules:

* Goal: be the last fighter alive.

* Turns: each round every fighter takes one turn. Turn order is randomized every round.

* Actions: Move one cell up/down/left/right, attack an adjacent enemy for 15–24 damage, or wait. An action consumes 1 AP.

* Rocks/Obstacles: 4 random impassable cells.

* Power-ups: Gold +1 AP per turn.

* Kills: the killer gets +1 AP per turn, and heals 50 HP (no over-heal).

(Unclear while watching replays, found in README.)

nananana93 days ago
This will be a weird rant, but the dialogue here is a perfect example of how SOTA models are so heavily tuned towards "solving agentic tasks" that they're useless at almost everything else - especially creative tasks.

"Coming for you, Crimson!"

"You'll never catch me alive, Azure!"

That's why nobody has been able to stick these things in a video game successfully, even though it seems like the tech is a perfect match.

It's all a game to them. They aren't afraid for their lives. They're making a mockery out of the world you've put them in. Those are not the words of little pixel people fighting to the death, those are AI abominations making "tool calls", LARPing as little pixel people fighting to the death.

I'm 100% serious when I say that you would've gotten cooler outputs with a GPT 3.5-era model, once you managed to beat it into producing structured output. Llama 2 would be giving the other agent a heartwarming story about how if it kills it there would be nobody to take care of its grandma or whatever, and the other agent would probably spare it.

The output is just so bland and devoid of soul. I feel like we would've found a lot of cool use cases for LLMs, had we not completely maimed their output in the pursuit of getting them to output 3% better TypeScript.

JarJarBeatU3 days ago
IMO this is because responding within a json response with all those other fields shifts the distribution to text that's staler.

If you told it to roleplay a pixel knight fighting to the death and provided it responses like "Authur (renamed claude sonnet 5) moves west towards you menacingly", the whole thing would come alive.

I've been working on an AI roleplaying game for a while now and it had felt stale until I separated the story output from the plumbing output (importantly, I do story first and then base the plumbing on that). I think the two just have incompatible distributions.

nananana93 days ago
I've been working on AI roleplaying too on and off, but after...

"Authur moves west towards you menacingly"

...RP is imemediately broken in the reasoning chain, and it's in that mode the prose is generated, e.g.

"Alright. Arthur moved west towards me menacingly. As Merlin, I should do X. No, I should do Y. Hmm. Maybe I should do Z. Let me reconsider. OK, I'll do X. How would Merlin respond? Maybe 'I'm a cool wizard'. No. 'I'm Merlin the wizard!' Yes, that's the load-bearing literary masteripiece. Let me respond with that!"

"I'm Merlin the wizard!"

bee_rider3 days ago
I’m not sure what good output in this case would look like. I’d expect “good prose” in a book’s fight scene to be action oriented and show something about the character or reflect something about the narrative, so, it should often not involve talking.

I guess realistic fights are just a lot of shouting and semi-coherent thoughts (because people don’t have time to come up with a full clever quip in that sort of situation).

mariofdistrust3 days ago
actually, it's mostly the harness' fault. messages have a character limit of only 50 and messages above that are silently trimmed. there's no space for anything interesting, though they still shouldn't have been that bad.
Lerc3 days ago
I think much of that is not due to limitations of their capability.

I think the text is bland because they are aiming to produce bland text.

Make it interesting in any particular direction and someone may not like it.

birdsongs3 days ago
> Make it interesting in any particular direction and someone may not like it.

The state of so much media and art these days. :/

isoprophlex3 days ago
The more advanced models clearly think ahead; they strategically wait for the others to mess eachother up, edging close to the battle but not so close they get caught up in the inital fighting. Then in most games they can finish off the survivors.

Makes you wonder, with enough intelligence and thinking budget, do they start to try to talk it out amongst eachother, staving off violence for longer and longer?

antoniojtorres3 days ago
Might have to add the battle royale storm/gas mechanic that closes in slowly.
vessenes3 days ago
Fun! As we know from 2024, these models can play diplomacy relatively well. Might I suggest they are allowed to talk to each other between rounds? Maybe up to two statements and emoji response to a DM. I think you’ll see more interesting gameplay.

Read the full thread on Hacker News →

Related stories