// NATURE NEWS — SPAZIO & SCIENZA
Scalable decision-making for games of imperfect information
Nature
volume 658, pages 55–59 (2026) Cite this article
Real-world decision-making generally involves hidden information, that is, information that is unknown to one agent but possessed by another. Unfortunately, the presence of large amounts of hidden information renders established reinforcement learning and search approaches ineffective. Even with multimillion-dollar industrial research efforts1, top-human-level play at Stratego—a board wargame with hidden information on a massive scale—has remained beyond the reach of artificial intelligence (AI). Here we introduce Ataraxos, an AI for Stratego based on general techniques that we developed for both self-play reinforcement learning and test-time search under hidden information. Ataraxos defeated the most decorated human Stratego player of all time by a large margin—achieving, to our knowledge, the first superhuman result in the game’s history—while consuming orders of magnitude less compute and data than previous efforts. Using the same techniques, we built a superhuman AI for Barrage Stratego and state-of-the-art AIs for Hanabi and dou dizhu, all with low cost and high sample efficiency. The success of this approach across adversarial, cooperative and team games establishes a design pattern for reinforcement learning and search that is effective under large amounts of hidden information, a longstanding desideratum of the field of strategic decision-making.
Imperfect information2,3 is a term of art describing interactions in which some agents may possess information that others do not. It is a characterizing feature of real-world settings, including financial markets, military conflict and negotiations.
Whereas in perfect-information settings, such as chess and Go, the conceptual ‘right thing to do’ is simple (select a move that maximizes the value of the resulting position), in those of imperfect information, it is subtle. The value of a decision depends not only on the policies agents employ thereafter but also on those they used—and counterfactually would have used—before and at the time of the decision (as these all affect the posterior distribution over hidden information). Even reasoning about the dependencies among contemporaneous counterfactuals in isolation is complex (Fig. 1).
The nodes represent decisions. The colour and line style of an edge represent whether changing the frequency of a decision on one end of the edge increases (blue and solid) or decreases (brown and dashed) the marginal value of the decision on the other end of the edge. The interconnectedness of frequencies and values across decisions makes it difficult to intuit the frequencies with which decisions should be made.
These dependencies have made developing artificial intelligence (AI) for imperfect-information settings challenging. The most successful approaches use sophisticated problem transformations based on public information4,5,6,7,8,9,10,11. But the cost of these transformations scales with the amount of hidden information, making them applicable only when this amount is small, as in Texas hold’em (in which there are 1,326 possible hands). Owing to this fundamental limitation, and the absence of an alternative foundation, strategic decision-making in settings with large amounts of hidden information has remained an open problem.
The unresolvedness of this problem is epitomized by the state of AI for Stratego—a board wargame resembling military chess (Fig. 2) prized as a challenge problem for the scale of its hidden information (there are over 1033 possible piece configurations). Stratego has been the subject of industrial research efforts spanning multiple years, involving dozens of researchers, and expending computation costing millions of dollars1. Yet, human players have remained superior, potentially making it the only classical game in which such well-resourced efforts have failed to produce superhuman performance.
a, The starting position of a game