Jev, an AI decision model, plays Pokémon Red all the way to the Hall of Fame. Watch the highlights.

281 points•pancomplex•6 days ago•124 comments•
Hey HN! Wanted to share a fun project I've been hacking on. Given Jev can make decisions really fast (but not fast enough to play Doom yet sadly), I wanted to try and push it to play a more complex game than Tetris. So I went with Pokémon.

I've spent endless hours playing this game as a child so building this was a ton of fun.

I open sourced everything in case you want to hack on it yourself here: https://github.com/christianmat/jev-pokemon

The game is being streamed live including the tokens and cost - hopefully we get all the badges and don't get stuck in a cave :)

124 comments

stusmall5 days ago
This is so interesting to watch. For a couple minutes I was in awe of how quick and cheap it was. Then I saw just how bad the decision are and how it would get stuck in strange loops of going in and out of the same door to no end.

This seems like a technology heading in the right direction but not quiet there yet. Excited for what they are cooking up but probably won't start building around it yet.

binlog5 days ago
This entire conversation around Jev seems weird to me. Like... we started from neural nets that could do basic decision making and classifications pretty well, then trained larger and larger language models to get to where we are now. Now suddenly everyone is going crazy because someone trained a smaller model that is adequate at making decisions? We already went through the "look this AI can play pokemon terribly" phase like a decade ago.
c7b5 days ago
A pre-trained universal classifier that can replace specifically-trained ones would have been considered just as much science fiction in the 2010's as the capabilities of modern LLMs. I'm not sure Jev is actually there yet, but at least it sounds theoretically doable today.

That being said, one thing having been unrealistic 10 years ago and just about possible today doesn't mean that it's going to change the world the same way another technically related, previously-impossible thing did. The Jev hype gives me a bit of the "you're still early to crypto" vibes of some later altcoins. I really like the idea, I think it's going to open up possibilities for using classifiers where we wouldn't or couldn't have trained one before. I'm crossing my fingers for an open weights version to drop. But it's still just a classifier, people have built similar things before Jev, the one thing that really stands out about it is their ability to generate hype.

solidasparagus5 days ago
The cheap, fast and smart-enough LLM space has been wildly neglected. Jev is one of the few players truly targeting that space. And for a lot of people it is the first time they are asking "what could I build if llms were interaction-speed fast?". The answers are cool, the problem is that Jev is not, I think, smart-enough yet to have that many applications, but it's smart enough that you can start to see what they will look like.
Garlef5 days ago
The difference is that you don't need training here; The decision graphs can be built on the fly by an LLM and contain instructions in plain language.

The magic moment for me from the Jev release was not that there was some system playing doom: Rather it was the moment, they just changed a part of the prompt to "don't shoot, just dodge" and the behavior changed immediately.

This means you can have a system with fast decision-making but still interact with it via language.

osener5 days ago
It is impressive, but all the hype and fake demos are selling it as a model that is as smart as frontier reasoning LLMs in the decisions it makes yet much cheaper and much faster, which is not true.
mtford5 days ago
This happens all the time in tech. A few years ago everybody got excited about static websites and server-side rendering as if we hadn't been doing that with PHP long ago.
ralusek5 days ago
The exact message I sent my friend this morning:

> the most interesting thing about this jev stuff

> is that people are seemingly like

> completely disinterested in how smart it actually is

> I haven't even heard it mentioned a single time how it actually compares to other LLMs coming up with their own classifications. Just: it's fast and cheap

After watching a few minutes of this it makes me think that maybe we should be a little more interested in how smart it is.

johnsmith18405 days ago
They know it's not. The company actually has or had public statements that they didn't like public benchmarks for comparison.

My main wonder is the difference between it and having a small llm no thinking output a single number only as a choice. Isn't that nearly the same here?

stusmall5 days ago
There is a lot of room for a lot of different models. For many use cases, intelligence beats out all.

For me in my day job, having extremely fast low quality decision makers over noisy inputs is very valuable. I work in security and having something that can help triage alerts, classify items and group things together is extremely valuable. It doesn't need to be perfect. Just being able to take a set of inputs from deterministic tooling and to be make general priority classifications goes a long way on helping humans look at the most important items first.

raincole5 days ago
Because it's not that smart especially when you compare it to other LLMs. The top LLMs have completely change the baseline of being smart.
jbjbjbjb5 days ago
Jev is for single shot classification, not multi-step RL environments with delayed reward and explore/exploit. My guess is it would go through the door with high confidence every time unless you change the input to add the history.
pancomplex5 days ago
This runs entirely on Jev as the only AI with a typescript harness that feeds it selective context.
Garlef5 days ago
I think the people behind jev made their intended use clear by labeling as "system one" - so they anticipate that there's a slower "system two" mediating jev
pancomplex5 days ago
Like others have mentioned in this post, I think a mix of models like Jev for simple stuff + a smarter reasoning model for more strategic thinking is the optimal solution. This experiment however is purely Jev. Which sometimes can be kinda dumb.
JamesSwift5 days ago
Well in a lot of ways this demo actually is a mix of jev for simple stuff + a very intelligent harness to fill in the rest. Pay attention to the "jev counter" on when api calls actually happen and the kinds of decisions its making. Its very rarely in a tight loop, and when it is (eg in an item menu) it tends to randomly walk through options.

Its very cool, but the harness is doing a _ton_ of heavy lifting.

stusmall5 days ago
Oh I want to be clear, I don't want my post to come off as negative. I really do mean things are headed the right direction and I think this harness is super cool. So don't take it that way and thank you for bringing something cool into this world.
mitxela5 days ago
feels like pre-LLM AI playing a video game but more expensive
ac2u5 days ago
Cool project, comes with a little too much guidance in the harness though IMO (pathfinding, textual milestones etc). (The author is very upfront about this in their README though)

I think if it was combined with a regular vLLM it could be really interesting, especially watching the reasoning logs.

Bonus points if it was one of the latest open models that somehow had all prior training knowledge of Pokemon abliterated so it was reasoning as an intelligent persona that had no knowledge of even the concept of Pokemon.

budoso5 days ago
I also tried to create a “Jev plays Pokémon” but with minimal additional context outside a move history and what is available in memory from the emulator, so no pathfinder, predetermined game path, etc. I can safely say that this experiment failed however, and Jev was not able to even get to Professor Oak’s lab to get a starter.

Props to OP for getting a working version, but it does not seem that this model is capable enough to play Pokémon at this point in time.

pancomplex4 days ago
Thank you. What I realized is that Jev is incredibly powerful for the decision making part of it when presented with a scenario that gives it a little context. It looks like we’re gonna be able to beat this game for under $2 of tokens. This is what blows my mind
MitPitt5 days ago
This is kinda chill to have in the background. I wish there were livestreams showing live reasoning of top models which are currently trying to solve cancer or whatever. Imagine the pogs in chat when it does.
pancomplex5 days ago
People need to be live streaming their AI more!
tehnoslow5 days ago
Actually, yes, that would be at least interesting
someothherguyy5 days ago
then it would cost human lives, no more fun i made this thing with no real effort vibes
dmitrygr5 days ago
Considering it just made Charizard forget its only fire-type move "Ember" to learn "Counter", I note no signs of intelligence.
pancomplex5 days ago
Rookie mistake clearly..
testaccount285 days ago
with such a fat harness, this is more like watching a walk thru play the game.
laszlokorte5 days ago
Yeah I would have expected it to only decide which button to press, not something abstract like the choice of "go east to lavender town" for the goal of "in lavender town, climb the pokemon tower"
pancomplex5 days ago
That was how I first implemented it. Jev sadly never left Pallet Town.
pancomplex4 days ago
Fair point, but I’d challenge anyone to build an ai or hard coded engine that, even following a walkthrough, completes the entire game for under $2 in tokens

Read the full thread on Hacker News →

Related stories