An AI system called Ataraxos has defeated four-time Stratego world champion Pim Niemeijer in a 20-game match, winning 15 games, losing one and drawing four. Researchers from Carnegie Mellon, MIT, New York University and Stanford trained the system with 16 GPUs for one week, plus four GPUs for four days for a supporting model.
Stratego is unusually difficult for machines because each player can see where 40 opposing pieces sit but not their identities. Games can last thousands of moves, and strong play requires bluffing while updating beliefs as pieces move and fight. The number of possible hidden arrangements makes exhaustive search impractical.
Ataraxos learned through 163 million self-play games. Its key addition was a second neural network that estimated plausible identities for hidden pieces from their movement. Before each move, the main system sampled likely arrangements, simulated candidate actions and selected a response based on those results.
The researchers estimate an earlier DeepMind system required 1,024 specialized chips for months, while Ataraxos used far less compute and played about 34 times fewer training games. The architecture also performed well in other imperfect-information games. Board games still provide fixed rules and clear winners, however, so success does not directly establish the system’s value in negotiations, markets or military planning.