A study published on September 30, 2026, reported that Ataraxos won 15 games, lost one and drew four against Pim Niemeijer in a 20-game Stratego evaluation. The series ran over three weeks. Ataraxos estimated the opponent’s hidden pieces and searched across plausible game states before choosing moves.

Ataraxos’s record against Pim Niemeijer

The researchers calculated an 85% effective win rate for the series by counting each draw as half a win. That figure describes this 20-game evaluation, not a general probability of winning.

The games used Strategus’s default 15+3 time control: each player had a 15-minute clock buffer and received three free seconds per move. Niemeijer was told Ataraxos would not adapt to his play during the series, while he could change his strategy between games. The study authors said the outcomes were not independent because human strategies changed across games.

How Ataraxos reasons about hidden pieces

Ataraxos combines self-play reinforcement learning with a belief network and search. Self-play lets the system train by playing games against itself. The belief network uses what a player can observe to estimate the types of the opponent’s hidden pieces; it does not reveal their actual identities.

For each candidate move, the search procedure samples plausible arrangements of those hidden pieces, runs game simulations from the sampled states and evaluates the resulting positions. This gives Ataraxos a way to weigh moves under uncertainty rather than treating one guessed board as certain.

A separate championship demonstration

At the Stratego World Championship demonstration on August 1–3, 2025, Ataraxos won 38 of 40 games, lost two and drew none. That event was separate from the 20-game evaluation against Niemeijer.

EvaluationPeriodOpponentsAtaraxos record
20-game evaluationThree weeksPim Niemeijer15 wins, 1 loss, 4 draws
World Championship demonstrationAugust 1–3, 2025Championship attendees38 wins, 2 losses, 0 draws

Training and other games

The Stratego training run used 16 NVIDIA H100 GPUs for one week. A separate run trained the belief network on four H100 GPUs for four days. Across its Stratego training, Ataraxos played 163 million completed self-play games and processed 208 billion environment steps.

The study authors also report results in other imperfect-information games. Ataraxos won four 50-game series in Barrage Stratego against three top-ranked players. The study reports a new state of the art in Hanabi and wins over earlier state-of-the-art bots in dou dizhu.