中文
相关论文

相关论文: Dynamics of Boltzmann Q-Learning in Two-Player Two…

200 篇论文

In this paper, we investigate the noncooperative games of multi-agent systems. Different from existing noncooperative games, our formulation involves the high-order nonlinear dynamics of players, and the communication topologies among…

系统与控制 · 电气工程与系统科学 2021-12-17 Zhenhua Deng , Jin Luo

We propose a novel independent and payoff-based learning framework for stochastic games that is model-free, game-agnostic, and gradient-free. The learning dynamics follow a best-response-type actor-critic architecture, where agents update…

机器学习 · 计算机科学 2026-02-03 Ahmed Said Donmez , Yuksel Arslantas , Muhammed O. Sayin

We examine the long-run behavior of multi-agent online learning in games that evolve over time. Specifically, we focus on a wide class of policies based on mirror descent, and we show that the induced sequence of play (a) converges to Nash…

计算机科学与博弈论 · 计算机科学 2022-08-11 Benoit Duvocelle , Panayotis Mertikopoulos , Mathias Staudigl , Dries Vermeulen

Zero-sum games are natural, if informal, analogues of closed physical systems where no energy/utility can enter or exit. This analogy can be extended even further if we consider zero-sum network (polymatrix) games where multiple agents…

计算机科学与博弈论 · 计算机科学 2019-03-06 James P. Bailey , Georgios Piliouras

We examine the tuning of cooperative behavior in repeated multi-agent games using an analytically tractable, continuous-time, nonlinear model of opinion dynamics. Each modeled agent updates its real-valued opinion about each available…

物理与社会 · 物理学 2021-11-24 Shinkyu Park , Anastasia Bizyaeva , Mari Kawakatsu , Alessio Franci , Naomi Ehrich Leonard

We consider evolutionary dynamics for population games in which players have a continuum of strategies at their disposal. Models in this setting amount to infinite-dimensional differential equations evolving on the manifold of probability…

动力系统 · 数学 2025-04-23 Brendon G. Anderson , Jingqi Li , Somayeh Sojoudi , Murat Arcak

Understanding the behavior of no-regret dynamics in general $N$-player games is a fundamental question in online learning and game theory. A folk result in the field states that, in finite games, the empirical frequency of play under…

Much of recent success in multiagent reinforcement learning has been in two-player zero-sum games. In these games, algorithms such as fictitious self-play and minimax tree search can converge to an approximate Nash equilibrium. While…

多智能体系统 · 计算机科学 2019-12-11 Alexander Shmakov , John Lanier , Stephen McAleer , Rohan Achar , Cristina Lopes , Pierre Baldi

This paper examines the convergence behaviour of simultaneous best-response dynamics in random potential games. We provide a theoretical result showing that, for two-player games with sufficiently many actions, the dynamics converge quickly…

计算机科学与博弈论 · 计算机科学 2025-05-19 Galit Ashkenazi-Golan , Domenico Mergoni Cecchelli , Edward Plumb

In this paper, we examine the long-run behavior of regularized, no-regret learning in finite games. A well-known result in the field states that the empirical frequencies of no-regret play converge to the game's set of coarse correlated…

计算机科学与博弈论 · 计算机科学 2023-11-07 Victor Boone , Panayotis Mertikopoulos

In large systems, it is important for agents to learn to act effectively, but sophisticated multi-agent learning algorithms generally do not scale. An alternative approach is to find restricted classes of games where simple, efficient…

多智能体系统 · 计算机科学 2009-03-16 Ian A. Kash , Eric J. Friedman , Joseph Y. Halpern

We discuss the long-run behavior of stochastic dynamics of many interacting players in spatial evolutionary games. In particular, we investigate the effect of the number of players and the noise level on the stochastic stability of Nash…

统计力学 · 物理学 2009-11-07 Jacek Miekisz

What happens when an infinite number of players play a quantum game? In this tutorial, we will answer this question by looking at the emergence of cooperation, in the presence of noise, in a one-shot quantum Prisoner's dilemma (QuPD). We…

物理与社会 · 物理学 2025-04-17 Colin Benjamin , Rajdeep Tah

In this article we evaluate the statistical evidence that a population of students learn about the sub-game perfect Nash equilibrium of the centipede game via repeated play of the game. This is done by formulating a model in which a…

统计方法学 · 统计学 2013-11-20 Anton H. Westveld , Peter D. Hoff

Finding the mixed Nash equilibria (MNE) of a two-player zero sum continuous game is an important and challenging problem in machine learning. A canonical algorithm to finding the MNE is the noisy gradient descent ascent method which in the…

最优化与控制 · 数学 2023-01-10 Yulong Lu

The approximation of mixed Nash equilibria (MNE) for zero-sum games with mean-field interacting players has recently raised much interest in machine learning. In this paper we propose a mean-field gradient descent dynamics for finding the…

最优化与控制 · 数学 2025-05-13 Yulong Lu , Pierre Monmarché

Cooperation is the foundation of ecosystems and the human society, and the reinforcement learning provides crucial insight into the mechanism for its emergence. However, most previous work has mostly focused on the self-organization at the…

物理与社会 · 物理学 2024-05-17 Zhen-Wei Ding , Guo-Zhong Zheng , Chao-Ran Cai , Wei-Ran Cai , Li Chen , Ji-Qiang Zhang , Xu-Ming Wang

We study two-player security games which can be viewed as sequences of nonzero-sum matrix games played by an Attacker and a Defender. The evolution of the game is based on a stochastic fictitious play process. Players do not have access to…

计算机科学与博弈论 · 计算机科学 2010-03-16 Kien C. Nguyen , Tansu Alpcan , Tamer Basar

Non-stationarity is a fundamental challenge in multi-agent reinforcement learning (MARL), where agents update their behaviour as they learn. Many theoretical advances in MARL avoid the challenge of non-stationarity by coordinating the…

计算机科学与博弈论 · 计算机科学 2025-03-19 Bora Yongacoglu , Gürdal Arslan , Serdar Yüksel

The theory of learning in games has extensively studied situations where agents respond dynamically to each other by optimizing a fixed utility function. However, in real situations, the strategic environment varies as a result of past…

计算机科学与博弈论 · 计算机科学 2022-07-15 Brandon C. Collins , Shouhuai Xu , Philip N. Brown