English

Superhuman AI for Stratego Using Self-Play Reinforcement Learning and Test-Time Search

Machine Learning 2025-11-11 v1 Artificial Intelligence

Abstract

Few classical games have been regarded as such significant benchmarks of artificial intelligence as to have justified training costs in the millions of dollars. Among these, Stratego -- a board wargame exemplifying the challenge of strategic decision making under massive amounts of hidden information -- stands apart as a case where such efforts failed to produce performance at the level of top humans. This work establishes a step change in both performance and cost for Stratego, showing that it is now possible not only to reach the level of top humans, but to achieve vastly superhuman level -- and that doing so requires not an industrial budget, but merely a few thousand dollars. We achieved this result by developing general approaches for self-play reinforcement learning and test-time search under imperfect information.

Keywords

Cite

@article{arxiv.2511.07312,
  title  = {Superhuman AI for Stratego Using Self-Play Reinforcement Learning and Test-Time Search},
  author = {Samuel Sokota and Eugene Vinitsky and Hengyuan Hu and J. Zico Kolter and Gabriele Farina},
  journal= {arXiv preprint arXiv:2511.07312},
  year   = {2025}
}
R2 v1 2026-07-01T07:30:13.348Z