Jev, an AI decision model, plays Pokémon Red all the way to the Hall of Fame. Watch the highlights.
I've spent endless hours playing this game as a child so building this was a ton of fun.
I open sourced everything in case you want to hack on it yourself here: https://github.com/christianmat/jev-pokemon
The game is being streamed live including the tokens and cost - hopefully we get all the badges and don't get stuck in a cave :)
124 comments
This seems like a technology heading in the right direction but not quiet there yet. Excited for what they are cooking up but probably won't start building around it yet.
That being said, one thing having been unrealistic 10 years ago and just about possible today doesn't mean that it's going to change the world the same way another technically related, previously-impossible thing did. The Jev hype gives me a bit of the "you're still early to crypto" vibes of some later altcoins. I really like the idea, I think it's going to open up possibilities for using classifiers where we wouldn't or couldn't have trained one before. I'm crossing my fingers for an open weights version to drop. But it's still just a classifier, people have built similar things before Jev, the one thing that really stands out about it is their ability to generate hype.
The magic moment for me from the Jev release was not that there was some system playing doom: Rather it was the moment, they just changed a part of the prompt to "don't shoot, just dodge" and the behavior changed immediately.
This means you can have a system with fast decision-making but still interact with it via language.
> the most interesting thing about this jev stuff
> is that people are seemingly like
> completely disinterested in how smart it actually is
> I haven't even heard it mentioned a single time how it actually compares to other LLMs coming up with their own classifications. Just: it's fast and cheap
After watching a few minutes of this it makes me think that maybe we should be a little more interested in how smart it is.
My main wonder is the difference between it and having a small llm no thinking output a single number only as a choice. Isn't that nearly the same here?
For me in my day job, having extremely fast low quality decision makers over noisy inputs is very valuable. I work in security and having something that can help triage alerts, classify items and group things together is extremely valuable. It doesn't need to be perfect. Just being able to take a set of inputs from deterministic tooling and to be make general priority classifications goes a long way on helping humans look at the most important items first.
Its very cool, but the harness is doing a _ton_ of heavy lifting.
I think if it was combined with a regular vLLM it could be really interesting, especially watching the reasoning logs.
Bonus points if it was one of the latest open models that somehow had all prior training knowledge of Pokemon abliterated so it was reasoning as an intelligent persona that had no knowledge of even the concept of Pokemon.
Props to OP for getting a working version, but it does not seem that this model is capable enough to play Pokémon at this point in time.
Read the full thread on Hacker News →
Related stories
- Hacker News · 3 points · 3 days ago
- Hacker News · 1 points · 2 days ago
- Hacker News · 421 points · 6 days ago
- Hacker News · 1 points · 6 days ago
- Hacker News · 1 points · 4 days ago
- Hacker News · 2 points · about 11 hours ago