English
Related papers

Related papers: Two-Timescale Q-Learning with Function Approximati…

200 papers

Computing approximate Nash equilibria in multi-player general-sum Markov games is a computationally intractable task. However, multi-player Markov games with certain cooperative or competitive structures might circumvent this…

Computer Science and Game Theory · Computer Science 2023-08-17 Zailin Ma , Jiansheng Yang , Zhihua Zhang

We investigate multi-agent reinforcement learning for stochastic games with complex tasks, where the reward functions are non-Markovian. We utilize reward machines to incorporate high-level knowledge of complex tasks. We develop an…

Multiagent Systems · Computer Science 2023-08-30 Jueming Hu , Jean-Raphael Gaglione , Yanze Wang , Zhe Xu , Ufuk Topcu , Yongming Liu

We derive the rate of convergence to Nash equilibria for the payoff-based algorithm proposed in \cite{tat_kam_TAC}. These rates are achieved under the standard assumption of convexity of the game, strong monotonicity and differentiability…

Optimization and Control · Mathematics 2022-02-24 Tatiana Tatarenko , Maryam Kamgarpour

Two-player complete-information game trees are perhaps the simplest possible setting for studying general-sum games and the computational problem of finding equilibria. These games admit a simple bottom-up algorithm for finding subgame…

Computer Science and Game Theory · Computer Science 2012-07-02 Michael L. Littman , Nishkam Ravi , Arjun Talwar , Martin Zinkevich

Two-time-scale Stochastic Approximation (SA) is an iterative algorithm with applications in reinforcement learning and optimization. Prior finite time analysis of such algorithms has focused on fixed point iterations with mappings…

Machine Learning · Computer Science 2025-09-30 Siddharth Chandak , Shaan Ul Haque , Nicholas Bambos

Nash equilibrium is a popular solution concept for solving imperfect-information games in practice. However, it has a major drawback: it does not preclude suboptimal play in branches of the game tree that are not reached in equilibrium.…

Computer Science and Game Theory · Computer Science 2017-05-29 Christian Kroer , Gabriele Farina , Tuomas Sandholm

We present a framework for computing approximate mixed-strategy Nash equilibria of continuous-action games. It is a modification of the traditional double oracle algorithm, extended to multiple players and continuous action spaces. Unlike…

Computer Science and Game Theory · Computer Science 2024-06-14 Carlos Martin , Tuomas Sandholm

Reinforcement learning from self-play has recently reported many successes. Self-play, where the agents compete with themselves, is often used to generate training data for iterative policy improvement. In previous work, heuristic rules are…

Machine Learning · Computer Science 2020-09-15 Yuanyi Zhong , Yuan Zhou , Jian Peng

As one of the most popular methods in the field of reinforcement learning, Q-learning has received increasing attention. Recently, there have been more theoretical works on the regret bound of algorithms that belong to the Q-learning class…

Machine Learning · Computer Science 2021-07-05 Zehao Dou , Zhuoran Yang , Zhaoran Wang , Simon S. Du

Using a probabilistic approach we study the parallel dynamics of fully connected Q-Ising neural networks for arbitrary Q. A Lyapunov function is shown to exist at zero temperature. A recursive scheme is set up to determine the time…

Disordered Systems and Neural Networks · Physics 2019-08-15 D. Bollé , G. Jongen , G. M. Shim

We study a finite-horizon two-person zero-sum risk-sensitive stochastic game for continuous-time Markov chains and Borel state and action spaces, in which payoff rates, transition rates and terminal reward functions are allowed to be…

Optimization and Control · Mathematics 2021-03-09 Junyu Zhang , Xianping Guo , Li Xia

This paper presents a pioneering investigation into discrete-time two-person non-zero-sum linear quadratic (LQ) stochastic games with random coefficients. We derive necessary and sufficient conditions for the existence of open-loop Nash…

Optimization and Control · Mathematics 2025-06-24 Yiwei Wu , Xun Li , Qingxin Meng

We consider payoff-based learning of a generalized Nash equilibrium (GNE) in multi-agent systems. Our focus is on games with jointly convex constraints of a linear structure and strongly monotone pseudo-gradients. We present a convergent…

Optimization and Control · Mathematics 2025-07-18 Tatiana Tatarenko , Maryam Kamgarpour

We motivate and propose a new model for non-cooperative Markov game which considers the interactions of risk-aware players. This model characterizes the time-consistent dynamic "risk" from both stochastic state transitions (inherent to the…

Computer Science and Game Theory · Computer Science 2019-11-22 Wenjie Huang , Pham Viet Hai , William B. Haskell

We give a converging semidefinite programming hierarchy of outer approximations for the set of quantum correlations of fixed dimension and derive analytical bounds on the convergence speed of the hierarchy. In particular, we give a…

Quantum Physics · Physics 2021-07-05 Hyejung H. Jee , Carlo Sparaciari , Omar Fawzi , Mario Berta

We introduce Q-Nash, a quantum annealing algorithm for the NP-complete problem of Fnding pure Nash equilibria in graphical games. The algorithm consists of two phases. The first phase determines all combinations of best response strategies…

Computer Science and Game Theory · Computer Science 2020-08-21 Christoph Roch , Thomy Phan , Sebastian Feld , Robert Müller , Thomas Gabor , Claudia Linnhoff-Popien

In this work, we present the first finite-time analysis of Q-learning with time-varying learning policies (i.e., on-policy sampling) for discounted Markov decision processes under minimal assumptions, requiring only the existence of a…

Machine Learning · Computer Science 2026-04-07 Phalguni Nanda , Zaiwei Chen

This paper investigates value function approximation in the context of zero-sum Markov games, which can be viewed as a generalization of the Markov decision process (MDP) framework to the two-agent case. We generalize error bounds from MDPs…

Artificial Intelligence · Computer Science 2013-01-07 Michail Lagoudakis , Ron Parr

In this paper, we propose a numerical methodology for finding the closed-loop Nash equilibrium of stochastic delay differential games through deep learning. These games are prevalent in finance and economics where multi-agent interaction…

Optimization and Control · Mathematics 2023-07-14 Robert Balkin , Hector D. Ceniceros , Ruimeng Hu

We present a new, distributed method to compute approximate Nash equilibria in bimatrix games. In contrast to previous approaches that analyze the two payoff matrices at the same time (for example, by solving a single LP that combines the…

Computer Science and Game Theory · Computer Science 2018-10-12 Artur Czumaj , Argyrios Deligkas , Michail Fasoulakis , John Fearnley , Marcin Jurdziński , Rahul Savani