There are difficult games, very difficult games, and then there's Montezuma's Revenge. This gem from the Atari 2600 (and other systems) is known for its absolute lack of mercy, with poor Panama Joe trapped in a nightmare of killer skulls, laser portals, snakes, spiders, falls, and death, lots of death. The folks at Google have spent a long time pitting their Deep-Q artificial intelligence against Montezuma's Revenge, and only after creating "artificial curiosity" did they manage to get it to explore a maximum of 15 screens, when before it didn't get past two.
A Personal Battle
Running back from school, ignoring homework, going to play soccer, coming back all dirty, my mom yelling from the kitchen, sitting in front of the TV, turning on the Atari 2600 clone I had back then… and turning Montezuma's Revenge into a personal challenge. Because it definitely "was" personal, no doubt. I think even today my parents regret buying "that machine." Montezuma's Revenge was one of the direct culprits in creating a real "graveyard of controllers," since I destroyed them jumping over skulls. Over time I got pretty good at the game… and a couple of decades later I played it on an emulator, only to rediscover an undeniable truth: I'm getting old and useless. The "8-year-old ninja reflexes" had completely disappeared. If Montezuma's Revenge was already unreachable for me, imagine an artificial intelligence trying to decipher it.
Deep-Q's Breakthrough
According to Google, their Deep-Q system had a rough time with Montezuma's Revenge. Last year we talked about its progress, and the high scores it achieved in fifty Atari 2600 games, but its performance in Montezuma's Revenge was null. Zero. Clearly something was wrong with the methodology and learning of the AI. Google's conclusion was that Deep-Q lacked curiosity, an extra stimulus to help it focus on exploration, and "try to win" the game as a human would. After several code tweaks and 100 million "training frames" later, Deep-Q only needed four tries to complete the first screen, when a year earlier it couldn't even score a single point.
With this new profile, Deep-Q managed to explore a total of fifteen screens out of the 24 that make up "level 1" of the game. In other words, it's still far from completing it, but Deep-Q's ambition aims much higher than we imagine. With each new version, the AI finds ways to polish its skills, and after what they did with AlphaGo, they're already talking about titles like StarCraft for its next phase. I figure someone in South Korea is eagerly awaiting that duel…