中文
相关论文

相关论文: Model-Free Learning for Two-Player Zero-Sum Partia…

200 篇论文

We introduce a stochastic learning process called the dampened gradient approximation process. While learning models have almost exclusively focused on finite games, in this paper we design a learning process for games with continuous…

计算机科学与博弈论 · 计算机科学 2018-07-02 Sebastian Bervoets , Mario Bravo , Mathieu Faure

We consider a general nonzero-sum impulse game with two players. The main mathematical contribution of the paper is a verification theorem which provides, under some regularity conditions, a suitable system of quasi-variational inequalities…

We address payoff-based decentralized learning in infinite-horizon zero-sum Markov games. In this setting, each player makes decisions based solely on received rewards, without observing the opponent's strategy or actions nor sharing…

计算机科学与博弈论 · 计算机科学 2025-02-11 Reda Ouhamma , Maryam Kamgarpour

Zero-sum stochastic games have found important applications in a variety of fields, from machine learning to economics. Work on this model has primarily focused on the computation of Nash equilibrium due to its effectiveness in solving…

计算机科学与博弈论 · 计算机科学 2022-11-28 Denizalp Goktas , Jiayi Zhao , Amy Greenwald

We study the sample complexity of identifying an approximate equilibrium for two-player zero-sum $n\times 2$ matrix games. That is, in a sequence of repeated game plays, how many rounds must the two players play before reaching an…

计算机科学与博弈论 · 计算机科学 2023-03-21 Arnab Maiti , Kevin Jamieson , Lillian J. Ratliff

Creating strong agents for games with more than two players is a major open problem in AI. Common approaches are based on approximating game-theoretic solution concepts such as Nash equilibrium, which have strong theoretical guarantees in…

计算机科学与博弈论 · 计算机科学 2018-11-07 Sam Ganzfried , Austin Nowak , Joannier Pinales

Imperfect recall games represent dynamic interactions where players forget previously known information, such as a history of played actions. The importance of imperfect recall games stems from allowing a concise representation of…

计算机科学与博弈论 · 计算机科学 2017-05-25 Jiri Cermak , Branislav Bosansky , Michal Pechoucek

Self-play is a technique for machine learning in multi-agent systems where a learning algorithm learns by interacting with copies of itself. Self-play is useful for generating large quantities of data for learning, but has the drawback that…

计算机科学与博弈论 · 计算机科学 2023-11-30 Revan MacQueen , James R. Wright

A recent body of experimental literature has studied empirical game-theoretical analysis, in which we have partial knowledge of a game, consisting of observations of a subset of the pure-strategy profiles and their associated payoffs to…

计算机科学与博弈论 · 计算机科学 2014-02-13 John Fearnley , Martin Gairing , Paul Goldberg , Rahul Savani

Computing the Nash equilibrium (NE) for N-player non-zerosum stochastic games is a formidable challenge. Currently, algorithmic methods in stochastic game theory are unable to compute NE for stochastic games (SGs) for settings in all but…

最优化与控制 · 数学 2021-03-25 David Mguni

Learning in zero-sum games studies a situation where multiple agents competitively learn their strategy. In such multi-agent learning, we often see that the strategies cycle around their optimum, i.e., Nash equilibrium. When a game…

计算机科学与博弈论 · 计算机科学 2025-03-06 Yuma Fujimoto , Kaito Ariu , Kenshi Abe

In this work, we study stochastic non-cooperative games, where only noisy black-box function evaluations are available to estimate the cost function for each player. Since each player's cost function depends on both its own decision…

计算机科学与博弈论 · 计算机科学 2025-11-18 Haidong Li , Anzhi Sheng , Yijie Peng , Long Wang

Computing approximate Nash equilibria in multi-player general-sum Markov games is a computationally intractable task. However, multi-player Markov games with certain cooperative or competitive structures might circumvent this…

计算机科学与博弈论 · 计算机科学 2023-08-17 Zailin Ma , Jiansheng Yang , Zhihua Zhang

In this paper, we consider multi-agent learning via online gradient descent in a class of games called $\lambda$-cocoercive games, a fairly broad class of games that admits many Nash equilibria and that properly includes unconstrained…

最优化与控制 · 数学 2021-07-20 Tianyi Lin , Zhengyuan Zhou , Panayotis Mertikopoulos , Michael I. Jordan

Multi-agent imitation learning (MA-IL) aims to learn optimal policies from expert demonstrations of interactions in multi-agent interactive domains. Despite existing guarantees on the performance of the resulting learned policies,…

机器学习 · 计算机科学 2026-02-25 Antoine Bergerault , Volkan Cevher , Negar Mehr

Consider a strongly monotone game where the players' utility functions include a reward function and a linear term for each dimension, with coefficients that are controlled by the manager. Gradient play converges to a unique Nash…

多智能体系统 · 计算机科学 2026-02-25 Siddharth Chandak , Ilai Bistritz , Nicholas Bambos

Unlike Poker where the action space $\mathcal{A}$ is discrete, differential games in the physical world often have continuous action spaces not amenable to discrete abstraction, rendering no-regret algorithms with…

计算机科学与博弈论 · 计算机科学 2025-02-17 Mukesh Ghimire , Zhe Xu , Yi Ren

Nash equilibrium (NE) is a widely adopted solution concept in game theory due to its stability property. However, we observe that the NE strategy might not always yield the best results, especially against opponents who do not adhere to NE…

人工智能 · 计算机科学 2024-08-13 Shuxin Li , Chang Yang , Youzhi Zhang , Pengdeng Li , Xinrun Wang , Xiao Huang , Hau Chan , Bo An

Computing Nash equilibrium (NE) of multi-player games has witnessed renewed interest due to recent advances in generative adversarial networks. However, computing equilibrium efficiently is challenging. To this end, we introduce the…

机器学习 · 计算机科学 2019-05-16 Arvind U. Raghunathan , Anoop Cherian , Devesh K. Jha

In this work, we propose novel offline and online Inverse Differential Game (IDG) methods for nonlinear Differential Games (DG), which identify the cost functions of all players from control and state trajectories constituting a feedback…

最优化与控制 · 数学 2024-11-18 Philipp Karg , Balint Varga , Sören Hohmann