English
Related papers

Related papers: Thompson Sampling Algorithm for Stochastic Games

200 papers

A strategy profile in a multi-player game is a Nash equilibrium if no player can unilaterally deviate to achieve a strictly better payoff. A profile is an $\epsilon$-Nash equilibrium if no player can gain more than $\epsilon$ by…

Computer Science and Game Theory · Computer Science 2026-01-27 Ali Asadi , Léonard Brice , Krishnendu Chatterjee , K. S. Thejaswini

This work considers stochastic differential games with a large number of players, whose costs and dynamics interact through the empirical distribution of both their states and their controls. We develop a new framework to prove convergence…

Probability · Mathematics 2022-03-24 Mathieu Laurière , Ludovic Tangpi

This work studies Nash equilibrium seeking for a class of stochastic aggregative games, where each player has an expectation-valued objective function depending on its local strategy and the aggregate of all players' strategies. We propose…

Optimization and Control · Mathematics 2022-05-17 Tongyu Wang , Peng Yi , Jie Chen

We show that Optimistic Hedge -- a common variant of multiplicative-weights-updates with recency bias -- attains ${\rm poly}(\log T)$ regret in multi-player general-sum games. In particular, when every player of the game uses Optimistic…

Machine Learning · Computer Science 2023-01-26 Constantinos Daskalakis , Maxwell Fishelson , Noah Golowich

We analyse the computational complexity of finding Nash equilibria in stochastic multiplayer games with $\omega$-regular objectives. While the existence of an equilibrium whose payoff falls into a certain interval may be undecidable, we…

Computer Science and Game Theory · Computer Science 2010-06-24 Michael Ummels , Dominik Wojtczak

We discuss a multiple-play multi-armed bandit (MAB) problem in which several arms are selected at each round. Recently, Thompson sampling (TS), a randomized algorithm with a Bayesian spirit, has attracted much attention for its empirically…

Machine Learning · Statistics 2019-03-22 Junpei Komiyama , Junya Honda , Hiroshi Nakagawa

This paper considers a time-varying game with $N$ players. Every time slot, players observe their own random events and then take a control action. The events and control actions affect the individual utilities earned by each player. The…

Computer Science and Game Theory · Computer Science 2014-02-04 Michael J. Neely

We revisit the Thompson sampling algorithm to control an unknown linear quadratic (LQ) system recently proposed by Ouyang et al (arXiv:1709.04047). The regret bound of the algorithm was derived under a technical assumption on the induced…

Systems and Control · Electrical Eng. & Systems 2022-09-21 Mukul Gagrani , Sagar Sudhakara , Aditya Mahajan , Ashutosh Nayyar , Yi Ouyang

In this paper, we treat linear quadratic team decision problems, where a team of agents minimizes a convex quadratic cost function over $T$ time steps subject to possibly distinct linear measurements of the state of nature. We assume that…

Optimization and Control · Mathematics 2022-12-23 Olle Kjellqvist , Ather Gattami

In this paper, we consider stochastic monotone Nash games where each player's strategy set is characterized by possibly a large number of explicit convex constraint inequalities. Notably, the functional constraints of each player may depend…

Optimization and Control · Mathematics 2023-08-25 Zeinab Alizadeh , Afrooz Jalilzadeh , Farzad Yousefian

We prove that in a normal form n-player game with m actions for each player, there exists an approximate Nash equilibrium where each player randomizes uniformly among a set of O(log(m) + log(n)) pure strategies. This result induces an…

Computer Science and Game Theory · Computer Science 2013-07-19 Yakov Babichenko , Ron Peretz

We consider learning Nash equilibria in two-player zero-sum Markov Games with nonlinear function approximation, where the action-value function is approximated by a function in a Reproducing Kernel Hilbert Space (RKHS). The key challenge is…

Machine Learning · Computer Science 2022-08-11 Chris Junchi Li , Dongruo Zhou , Quanquan Gu , Michael I. Jordan

This paper introduces a new method to achieve stable convergence to Nash equilibrium in duopoly noncooperative games. Inspired by the recent fixed-time Nash Equilibrium seeking (NES) as well as prescribed-time extremum seeking (ES) and…

Optimization and Control · Mathematics 2024-05-27 Victor Hugo Pereira Rodrigues , Tiago Roux Oliveira , Miroslav Krstić , Tamer Başar

Learning in games refers to scenarios where multiple players interact in a shared environment, each aiming to minimize their regret. An equilibrium can be computed at a fast rate of $O(1/T)$ when all players follow the optimistic…

Computer Science and Game Theory · Computer Science 2025-02-18 Taira Tsuchiya , Shinji Ito , Haipeng Luo

We develop a form Thompson sampling for online learning under full feedback - also known as prediction with expert advice - where the learner's prior is defined over the space of an adversary's future actions, rather than the space of…

Machine Learning · Computer Science 2025-09-23 Alexander Terenin , Jeffrey Negrea

We study two-player zero-sum stochastic games, and propose a form of independent learning dynamics called Doubly Smoothed Best-Response dynamics, which integrates a discrete and doubly smoothed variant of the best-response dynamics into…

Computer Science and Game Theory · Computer Science 2023-03-07 Zaiwei Chen , Kaiqing Zhang , Eric Mazumdar , Asuman Ozdaglar , Adam Wierman

This work considers a stochastic Nash game in which each player solves a parameterized stochastic optimization problem. In deterministic regimes, best-response schemes have been shown to be convergent under a suitable spectral property…

Optimization and Control · Mathematics 2018-02-08 Jinlong Lei , Uday V. Shanbhag , Jong-Shi Pang , Suvrajeet Sen

The empirically successful Thompson Sampling algorithm for stochastic bandits has drawn much interest in understanding its theoretical properties. One important benefit of the algorithm is that it allows domain knowledge to be conveniently…

Machine Learning · Computer Science 2016-07-22 Che-Yu Liu , Lihong Li

We study the existence of mixed-strategy equilibria in concurrent games played on graphs. While existence is guaranteed with safety objectives for each player, Nash equilibria need not exist when players are given arbitrary terminal-reward…

Computer Science and Game Theory · Computer Science 2016-09-15 Patricia Bouyer , Nicolas Markey , Daniel Stan

We propose a model of inter-bank lending and borrowing which takes into account clearing debt obligations. The evolution of log-monetary reserves of $N$ banks is described by coupled diffusions driven by controls with delay in their drifts.…

Mathematical Finance · Quantitative Finance 2016-07-22 Rene Carmona , Jean-Pierre Fouque , Seyyed Mostafa Mousavi , Li-Hsien Sun
‹ Prev 1 4 5 6 7 8 10 Next ›