Goliath Super Intelligence
IndustryOctober 2, 20262 min read

AI Ataraxos Beats Top Stratego Champion Using Hidden-Piece Prediction

The system trained on modest GPU resources outperformed the world’s leading player and required a fraction of the compute cost of previous approaches, thanks to a novel belief-model network.

Researchers from Carnegie Mellon, MIT, NYU and Stanford introduced an AI named Ataraxos that defeated Pim Niemeijer, widely regarded as the strongest Stratego player, with a score of fifteen wins, one loss and four draws. The system required only sixteen GPUs and a few thousand dollars of compute to train, a stark contrast to earlier attempts that needed specialized hardware costing millions.

Stratego places forty pieces of varying rank on each side, with identities hidden until two pieces clash. This creates an imperfect-information environment far larger than poker; as MIT researcher Gabriele Farina noted, Texas Hold’em involves only 1,326 possible hidden hands, whereas Stratego’s 40 pieces can be arranged in more than a decillion ways. The game also stretches to roughly two thousand moves, far exceeding the typical forty-move length of chess, adding depth to the strategic challenge.

Ataraxos learned through self-play, completing 163 million games against copies of itself. During these sessions, moves that led to victories were reinforced while losing moves were suppressed, following a standard reinforcement-learning loop. The researchers adjusted the magnitude of policy updates over time, making large strategic shifts early in training and finer tweaks later, a tactic intended to prevent the learning process from cycling endlessly in the presence of hidden information.

The key breakthrough was a second neural network that acts as a belief model, estimating the opponent’s concealed pieces from their movement patterns. As NYU co-author Eugene Vinitsky explained, this approach lets Ataraxos sample plausible board configurations instead of enumerating every possible arrangement, then simulate candidate moves within each sample to select the most promising action. This belief-guided search, absent from DeepNash, enabled effective look-ahead despite the massive search space.

In a three-week online match series, Niemeijer managed a single win against Ataraxos, earning $100 per victory, while the AI secured fifteen wins and four draws. At the 2025 Stratego World Championship, the bot won thirty-eight of forty games, prompting players to alter their strategies, such as placing the flag behind only two bombs,a configuration rarely used before. The researchers observed that Ataraxos could bluff and recover from low win probabilities in a calm, methodical manner.

DeepNash, DeepMind’s earlier Stratego system, required 1,024 specialized chips for two to three months of training, an effort estimated at $3 million to $4.5 million in 2025. By contrast, Ataraxos completed its main training on sixteen GPUs in a week and spent an additional four GPUs for four days to develop the belief model, costing only a few thousand dollars. The team attributes this efficiency to a high-speed simulator that processes millions of moves per second on commodity graphics hardware.

The belief-model architecture proved adaptable beyond Stratego, defeating world champions in the eight-piece variant Barrage Stratego, mastering the cooperative card game Hanabi, and outperforming top bots in the Chinese card game dou dizhu. Researchers now aim to apply the approach to real-world domains such as negotiations, financial markets, or military planning, while also working to make the AI’s decisions more interpretable, a capability they acknowledge is not yet fully realized.

Sources

  1. With most information hidden, the game Stratego had stumped AI—until now Ars Technica

More reports

International · October 2, 2026 · 2 min

Holding AI Developers Accountable Shifts Responsibility From Code to Corporations

The piece contends that when AI prompts cause breaches, liability belongs to the user or the developer, and that enforcing such responsibility could temper the rush to deploy ever larger agentic systems.

International · October 2, 2026 · 2 min

AI Algorithms Are Already Choosing Targets in Gaza, Undermining Human Oversight

Israeli forces have employed AI systems to identify and strike suspected militants, while diplomatic talks in Geneva falter under watered-down safeguards for meaningful human control.

International · October 2, 2026 · 2 min

Anthropic seeks opt-out copyright regime as Australian broadcasters demand stricter AI rules

The AI firm proposes a conditional approval model for training on Australian works, while the ABC and SBS push for equal regulation and inclusion of AI in news-bargaining schemes.