Developer Monty Bichouna reported that a world model selected a starter in 52 of 100 plans in Pokémon Red. The planner began from a fixed save in Professor Oak’s lab, and its goal was to choose any starter—not to play through the game. Bichouna published the project account on September 20, 2026.

How the model predicted what button presses would do

Bichouna built the model around a Joint-Embedding Predictive Architecture, or JEPA, approach. Rather than generate a full screenshot for every imagined move, it encoded a game frame into a compressed, 192-number representation—called a latent embedding—and predicted how that representation would change after an action.

The reported training set contained 42,382 grayscale frames across 1,009 short trajectories. The routes included scripted play, noisy button inputs and more random movement, giving the model examples of different ways the game could change after an action.

Training focused on predicting those changes: it did not use a reward for acquiring a Pokémon. The starter goals entered later, when the planner searched for button sequences predicted to reach them.

The 100-plan result

Bichouna reported testing 100 plans with the model and starting save held fixed, using fresh random seeds. The same evaluation account gives the outcomes for random button sequences and for the search procedure with an untrained predictor:

MethodSuccessful starter selectionsPlans tested
Trained planner52100
Random button sequences0100
Search with an untrained predictor1100

Within this task, the trained planner’s result stands apart from the two comparisons. The untrained predictor barely changed the outcome relative to random button presses, while the trained model selected a starter in 52 plans.

The planner’s target was a short sequence, not a playthrough

Each candidate plan contained 14 button presses. In each search round, the planner sampled 512 candidate plans, kept the 64 with the lowest predicted distance to a starter goal, then updated how it sampled the next round.

That setup began from a save already inside Professor Oak’s lab. Any of the three starter Pokémon counted as success. In the documented successful plan, the model selected Squirtle; the emulator’s party count went from zero to one. The experiment stopped at starter selection.

Model size, hardware and code

Bichouna reported that the final model had about 12.5 million parameters and was trained locally on one Nvidia GeForce RTX 3080 Ti. The developer’s open-source implementation is called lePokeRed.