中文
相关论文

相关论文: Improving Sample Efficiency of Model-Free Algorith…

200 篇论文

We contribute the first provable guarantees of global convergence to Nash equilibria (NE) in two-player zero-sum convex Markov games (cMGs) by using independent policy gradient methods. Convex Markov games, recently defined by Gemp et al.…

计算机科学与博弈论 · 计算机科学 2025-06-23 Fivos Kalogiannis , Emmanouil-Vasileios Vlatakis-Gkaragkounis , Ian Gemp , Georgios Piliouras

Markov games provide a powerful framework for modeling strategic multi-agent interactions in dynamic environments. Traditionally, convergence properties of decentralized learning algorithms in these settings have been established only for…

多智能体系统 · 计算机科学 2025-06-13 Chinmay Maheshwari , Manxi Wu , Shankar Sastry

Decision-making in multi-player games can be extremely challenging, particularly under uncertainty. In this work, we propose a new sample-based approximation to a class of stochastic, general-sum, pure Nash games, where each player has an…

We study best-response type learning dynamics for zero-sum polymatrix games under two information settings. The two settings are distinguished by the type of information that each player has about the game and their opponents' strategy. The…

最优化与控制 · 数学 2025-08-13 Fathima Zarin Faizal , Asuman Ozdaglar , Martin J. Wainwright

Regret minimization is a general approach to online optimization which plays a crucial role in many algorithms for approximating Nash equilibria in two-player zero-sum games. The literature mainly focuses on solving individual games in…

计算机科学与博弈论 · 计算机科学 2025-04-29 David Sychrovský , Martin Schmid , Michal Šustr , Michael Bowling

This research paper introduces a model-free optimal controller for discrete-time Markovian jump linear systems (MJLSs), employing principles from the methodology of reinforcement learning (RL). While Q-learning methods have demonstrated…

系统与控制 · 电气工程与系统科学 2024-08-07 Ehsan Badfar , Babak Tavassoli

This paper initiates the study of scale-free learning in Markov Decision Processes (MDPs), where the scale of rewards/losses is unknown to the learner. We design a generic algorithmic framework, \underline{S}cale \underline{C}lipping…

机器学习 · 计算机科学 2024-03-05 Mingyu Chen , Xuezhou Zhang

As quantum processors advance, the emergence of large-scale decentralized systems involving interacting quantum-enabled agents is on the horizon. Recent research efforts have explored quantum versions of Nash and correlated equilibria as…

计算机科学与博弈论 · 计算机科学 2024-12-18 Wayne Lin , Georgios Piliouras , Ryann Sim , Antonios Varvitsiotis

A fundamental question in reinforcement learning is whether model-free algorithms are sample efficient. Recently, Jin et al. \cite{jin2018q} proposed a Q-learning algorithm with UCB exploration policy, and proved it has nearly optimal…

机器学习 · 计算机科学 2019-09-30 Kefan Dong , Yuanhao Wang , Xiaoyu Chen , Liwei Wang

We address in this paper Reinforcement Learning (RL) among agents that are grouped into teams such that there is cooperation within each team but general-sum (non-zero sum) competition across different teams. To develop an RL method that…

机器学习 · 计算机科学 2025-02-11 Muhammad Aneeq uz Zaman , Alec Koppel , Mathieu Laurière , Tamer Başar

In this paper, we investigate a competitive market involving two agents who consider both their own wealth and the wealth gap with their opponent. Both agents can invest in a financial market consisting of a risk-free asset and a risky…

最优化与控制 · 数学 2025-02-10 Junyi Guo , Xia Han , Hao Wang , Kam Chuen Yuen

There are only a few learning algorithms applicable to stochastic dynamic teams and games which generalize Markov decision processes to decentralized stochastic control problems involving possibly self-interested decision makers. Learning…

最优化与控制 · 数学 2016-05-03 Gürdal Arslan , Serdar Yüksel

Mean-field games have been used as a theoretical tool to obtain an approximate Nash equilibrium for symmetric and anonymous $N$-player games. However, limiting applicability, existing theoretical results assume variations of a "population…

最优化与控制 · 数学 2023-06-12 Batuhan Yardim , Semih Cayci , Matthieu Geist , Niao He

We consider the problem of simultaneous learning in stochastic games with many players in the finite-horizon setting. While the typical target solution for a stochastic game is a Nash equilibrium, this is intractable with many players. We…

计算机科学与博弈论 · 计算机科学 2022-10-27 William Brown

Reinforcement learning (RL) so far has limited real-world applications. One key challenge is that typical RL algorithms heavily rely on a reset mechanism to sample proper initial states; these reset mechanisms, in practice, are expensive to…

机器学习 · 计算机科学 2023-07-25 Hoai-An Nguyen , Ching-An Cheng

We present the convergence rates of synchronous and asynchronous Q-learning for average-reward Markov decision processes, where the absence of contraction poses a fundamental challenge. Existing non-asymptotic results overcome this…

机器学习 · 计算机科学 2026-01-30 Zijun Chen , Zaiwei Chen , Nian Si , Shengbo Wang

Multi-agent imitation learning (MA-IL) aims to learn optimal policies from expert demonstrations of interactions in multi-agent interactive domains. Despite existing guarantees on the performance of the resulting learned policies,…

机器学习 · 计算机科学 2026-02-25 Antoine Bergerault , Volkan Cevher , Negar Mehr

We address learning Nash equilibria in convex games under the payoff information setting. We consider the case in which the game pseudo-gradient is monotone but not necessarily strictly monotone. This relaxation of strict monotonicity…

最优化与控制 · 数学 2023-08-17 Tatiana Tatarenko , Maryam Kamgarpour

We establish an optimal sample complexity of $O(\epsilon^{-2})$ for obtaining an $\epsilon$-optimal global policy using a single-timescale actor-critic (AC) algorithm in infinite-horizon discounted Markov decision processes (MDPs) with…

机器学习 · 计算机科学 2026-05-08 Navdeep Kumar , Tehila Dahan , Lior Cohen , Ananyabrata Barua , Giorgia Ramponi , Kfir Yehuda Levy , Shie Mannor

We study distributionally robust Markov games (DR-MGs) with the average-reward criterion, a framework for multi-agent decision-making under uncertainty over extended horizons. In average reward DR-MGs, agents aim to maximize their…

多智能体系统 · 计算机科学 2025-12-12 Zachary Roch , Yue Wang
‹ 上一页 1 8 9 10 下一页 ›