通过无模型多智能体强化学习掌握 Stratego 游戏
人工智能
2023-01-11 v1 计算机科学与博弈论
多智能体系统
摘要
我们提出 DeepNash,一种能够从零开始学习玩不完全信息游戏 Stratego 直至人类专家水平的自主智能体。Stratego 是人工智能(AI)尚未攻克的少数标志性棋盘游戏之一。这款流行游戏的博弈树规模约为 个节点,即比围棋大 倍。它还具有在不完全信息下决策额外的复杂性,类似于德州扑克,而后者的博弈树显著更小(约为 个节点)。Stratego 中的决策是在大量离散动作上做出的,动作与结果之间无明显联系。对局很长,通常一方获胜前需数百步,且 Stratego 中的局面不能像扑克那样轻易分解为可处理的子问题。由于这些原因,Stratego 数十年来一直是 AI 领域的重大挑战,现有 AI 方法仅达到业余水平。DeepNash 使用一种博弈论的无模型深度强化学习方法,无需搜索,通过自博弈学习掌握 Stratego。正则化纳什动态(Regularised Nash Dynamics, R-NaD)算法作为 DeepNash 的关键组件,通过直接修改底层多智能体学习动态,收敛到近似纳什均衡而非围绕其“循环”。DeepNash 在 Stratego 中击败现有最先进(state-of-the-art, SOTA)AI 方法,并在 Gravon 游戏平台取得年度(2022)及历史前三排名,与人类专家玩家同台竞技。
引用
@article{arxiv.2206.15378,
title = {Mastering the Game of Stratego with Model-Free Multiagent Reinforcement Learning},
author = {Julien Perolat and Bart de Vylder and Daniel Hennes and Eugene Tarassov and Florian Strub and Vincent de Boer and Paul Muller and Jerome T. Connor and Neil Burch and Thomas Anthony and Stephen McAleer and Romuald Elie and Sarah H. Cen and Zhe Wang and Audrunas Gruslys and Aleksandra Malysheva and Mina Khan and Sherjil Ozair and Finbarr Timbers and Toby Pohlen and Tom Eccles and Mark Rowland and Marc Lanctot and Jean-Baptiste Lespiau and Bilal Piot and Shayegan Omidshafiei and Edward Lockhart and Laurent Sifre and Nathalie Beauguerlange and Remi Munos and David Silver and Satinder Singh and Demis Hassabis and Karl Tuyls},
journal= {arXiv preprint arXiv:2206.15378},
year = {2023}
}