PacBench: how well a model and harness can recreate Pac-Man from a single prompt.

78 points•thefourthchime•2 days ago•50 comments•
Benchmarks how well Harness+models can create a Pac-Man game from a single prompt:

“Create a Pac-Man game in a single HTML page”

Each model gets one shot — no follow-up prompts or fixes.

50 comments

yambam1 day ago
This reminds me of a couple decades ago when some friends and I set ourselves a challenge to, individually, each create as much of a Pac-Man clone as possible in 10 hours. None of us had any experience in games programming or graphics coding. It was great fun and we all learned a lot.

No-one ended up with a complete clone but I loved how we all ended up focusing on different things, like pixel-perfect graphics versus accuracy in gameplay, and how we all brought our existing skills to the challenge despite not really knowing what we were doing.

I expect if we had AI models available it would have ruined the pleasure of figuring it out for ourselves. I feel kind of sad for the next generation of developers who won't have that experience.

OJFord1 day ago
I think you figure out different things for yourself; the next generation is going to have 'programming' appeal to them for different reasons. More creatives. More engineers, perhaps, as the practice becomes more about the bigger picture.
HPsquared1 day ago
It'll be more like architecture than engineering, I think. Sketch out the building plan, context (physical and social) and desired usage.
drcxd1 day ago
Interesting, recently I am working on my own clone of Pac-Man. LLM implementations lose lots of details. They are not 1:1 replication of the original game. For example, the behavior of the ghost is not the same as the original. If I have not implemented the game myself, I can not tell the differences. What LLM produced look like the original game, but they are not.
CamperBob21 day ago
But it takes -- what, two or three sentences? -- to explain how the ghosts should move. The interesting (and important) thing is not that the model gets it wrong at first, it's how easy it is to correct it.

One-shotting something like Pac-Man doesn't prove much. At the end of the day, one-shot fidelity is going to scale more or less linearly with model size/world knowledge. Why wouldn't it?

strataspace2 days ago
I tried this with DOOM. Fable 5 did a pretty shit job. Astra made pretty crazy animated sprites and was pretty good considering.

The fact that these are at all playable and 100x my programming skill level is pretty depressing from a certain pov. The ThreeJS dude posted ab how demotivated he was to continue his work, and while I was never a dev that did much with webgl, I commiserate.

acomjean2 days ago
I wonder if we need better programming abstractions/ languages that can make programming easier for people.

It seems like a lot of the programming is cookie cutter type stuff (where these ai programs shine) and should be easier.

They can be amazing, but promoting “make a pac man game” seems like an exercise in which model has the best training code that was Pac-Man.

Joel_Mckay1 day ago
>I wonder if we need better programming abstractions/ languages that can make programming easier for people.

That was nodejs, and it turns out people unwilling and or unable to improve things through systematic skill refinement just degrades into chaos.

Successful ecosystems are just pedantically lame enough to keep silly folks from going YOLO, but empowering enough to allow people to still have fun.

>which model has the best training code that was Pac-Man.

You mean which model is more cautious about copyright bleed-through of the $9Tn in FOSS and user code they misappropriated though isomorphic plagiarism. =3

shoobiedoo2 days ago
My hope is that this drives a new generation of hyper creative content from those who don't give up. I mean, of course it does pacman well. Pacman and its clones have been done to the death over decades. But what if we ask AI to write finnegan's wake 2?
Vakaiser2 days ago
I’m quite optimistic about how AI tooling will raise the floor for creative work.

I personally have been working on a Three.js project with Opus 5 and 5.5 that I never would have continued with had I needed to dive into documentation by hand.

Seeing immediate results is incredibly motivating.

fuzzythinker1 day ago
Hope this gives a speck of encouragement. Just want to call out how I love the clarity in Three.js Resources, in readability, design, and answering why. Eg. https://threejsresources.com/tsl
lukan1 day ago
"The ThreeJS dude posted ab how demotivated he was to continue his work"

Mrdoob? If so, can you show a link?

jmathai2 days ago
This prompt is a good way to test how well models fill in missing context because it's so nondescript. They're definitely improving.

Remember when people considered you a genius for prompting with "You are a skilled writer....".

wredcoll2 days ago
That moment in time was physically painful.
_matthew_2 days ago
I don't think it makes sense to have the prompt be that short. This is basically a bench.ark of how models interpret an overly vague prompt. It should at least be "Create a pacman clone in a single html page. Make it faithful to the original" if that's what we're scoring it on.

Read the full thread on Hacker News →

Related stories