中文
相关论文

相关论文: On Reinforcement Learning for Turn-based Zero-sum …

200 篇论文

This paper resolves the open question of designing near-optimal algorithms for learning imperfect-information extensive-form games from bandit feedback. We present the first line of algorithms that require only…

机器学习 · 计算机科学 2023-04-04 Yu Bai , Chi Jin , Song Mei , Tiancheng Yu

The design of Nash equilibrium seeking strategies for games in which the involved players are of second-order integrator-type dynamics is investigated in this paper. Noticing that velocity signals are usually noisy or not available for…

最优化与控制 · 数学 2020-06-18 Maojiao Ye , Jizhao Yin , Le Yin

We consider the problem of computing mixed Nash equilibria of two-player zero-sum games with continuous sets of pure strategies and with first-order access to the payoff function. This problem arises for example in game-theory-inspired…

最优化与控制 · 数学 2025-09-04 Guillaume Wang , Lénaïc Chizat

Solving Nash equilibrium is the key challenge in normal-form games with large strategy spaces, where open-ended learning frameworks offer an efficient approach. In this work, we propose an innovative unified open-ended learning framework…

计算机科学与博弈论 · 计算机科学 2024-03-25 Yudong Hu , Haoran Li , Congying Han , Tiande Guo , Mingqiang Li , Bonan Li

Many real-world applications can be described as large-scale games of imperfect information. To deal with these challenging domains, prior work has focused on computing Nash equilibria in a handcrafted abstraction of the domain. In this…

机器学习 · 计算机科学 2016-06-29 Johannes Heinrich , David Silver

The nested Extremum Seeking (nES) algorithm is a model-free optimization method that has been shown to converge to a neighborhood of a Nash equilibrium. In this work, we demonstrate that the same nES dynamics can instead be made to converge…

最优化与控制 · 数学 2026-04-02 Brad Ratto , Alan Williams , Miroslav Krstić , Tamer Başar , Alexander Scheinker

We study the game modification problem, where a benevolent game designer or a malevolent adversary modifies the reward function of a zero-sum Markov game so that a target deterministic or stochastic policy profile becomes the unique Markov…

计算机科学与博弈论 · 计算机科学 2024-08-27 Young Wu , Jeremy McMahan , Yiding Chen , Yudong Chen , Xiaojin Zhu , Qiaomin Xie

We initiate the study of how to perturb the reward in a zero-sum Markov game with two players to induce a desirable Nash equilibrium, namely arbitrating. Such a problem admits a bi-level optimization formulation. The lower level requires…

多智能体系统 · 计算机科学 2023-02-21 Jing Wang , Meichen Song , Feng Gao , Boyi Liu , Zhaoran Wang , Yi Wu

One key in real-life Nash equilibrium applications is to calibrate players' cost functions. To leverage the approximation ability of neural networks, we proposed a general framework for optimizing and learning Nash equilibrium using neural…

计算机科学与博弈论 · 计算机科学 2024-09-04 Di Zhang , Wei Gu , Qing Jin

In this work, we adapt a training approach inspired by the original AlphaGo system to play the imperfect information game of Reconnaissance Blind Chess. Using only the observations instead of a full description of the game state, we first…

人工智能 · 计算机科学 2022-08-04 Timo Bertram , Johannes Fürnkranz , Martin Müller

Nash equilibrium is a key concept in game theory fundamental for elucidating the equilibrium state of strategic interactions, finding applications in diverse fields such as economics, political science, and biology. However, the Nash…

计算机科学与博弈论 · 计算机科学 2024-04-02 Elie Eshoa , Ali R. Zomorrodi

Nash equilibrium is a central concept in game theory. Several Nash solvers exist, yet none scale to normal-form games with many actions and many players, especially those with payoff tensors too big to be stored in memory. In this work, we…

计算机科学与博弈论 · 计算机科学 2022-02-07 Ian Gemp , Rahul Savani , Marc Lanctot , Yoram Bachrach , Thomas Anthony , Richard Everett , Andrea Tacchetti , Tom Eccles , János Kramár

This paper aims to accommodate games in which the players' dynamics are subject to un-modeled and disturbance terms. The un-modeled and disturbance terms are regarded as extended states for which observers are designed to estimate them.…

最优化与控制 · 数学 2020-04-22 Maojiao Ye

This article discusses two contributions to decision-making in complex partially observable stochastic games. First, we apply two state-of-the-art search techniques that use Monte-Carlo sampling to the task of approximating a…

计算机科学与博弈论 · 计算机科学 2014-01-21 Marc Ponsen , Steven de Jong , Marc Lanctot

Much of recent success in multiagent reinforcement learning has been in two-player zero-sum games. In these games, algorithms such as fictitious self-play and minimax tree search can converge to an approximate Nash equilibrium. While…

多智能体系统 · 计算机科学 2019-12-11 Alexander Shmakov , John Lanier , Stephen McAleer , Rohan Achar , Cristina Lopes , Pierre Baldi

We propose a reinforcement learning algorithm for stationary mean-field games, where the goal is to learn a pair of mean-field state and stationary policy that constitutes the Nash equilibrium. When viewing the mean-field state and the…

机器学习 · 计算机科学 2020-10-12 Qiaomin Xie , Zhuoran Yang , Zhaoran Wang , Andreea Minca

Constrained Markov games offer a formal mathematical framework for modeling multi-agent reinforcement learning problems where the behavior of the agents is subject to constraints. In this work, we focus on the recently introduced class of…

机器学习 · 计算机科学 2024-02-29 Philip Jordan , Anas Barakat , Niao He

Adversarial team games model multiplayer strategic interactions in which a team of identically-interested players is competing against an adversarial player in a zero-sum game. Such games capture many well-studied settings in game theory,…

Finding approximate Nash equilibria in zero-sum imperfect-information games is challenging when the number of information states is large. Policy Space Response Oracles (PSRO) is a deep reinforcement learning algorithm grounded in game…

计算机科学与博弈论 · 计算机科学 2021-02-22 Stephen McAleer , John Lanier , Roy Fox , Pierre Baldi

Adversarial training is a standard technique for training adversarially robust models. In this paper, we study adversarial training as an alternating best-response strategy in a 2-player zero-sum game. We prove that even in a simple…

机器学习 · 计算机科学 2023-03-01 Maria-Florina Balcan , Rattana Pukdee , Pradeep Ravikumar , Hongyang Zhang