English
Related papers

Related papers: Thompson Sampling Algorithm for Stochastic Games

200 papers

In an $\epsilon$-Nash equilibrium, a player can gain at most $\epsilon$ by unilaterally changing his behaviour. For two-player (bimatrix) games with payoffs in $[0,1]$, the best-known$\epsilon$ achievable in polynomial time is 0.3393. In…

Computer Science and Game Theory · Computer Science 2014-10-02 Argyrios Deligkas , John Fearnley , Rahul Savani , Paul Spirakis

This paper investigates stochastic generalized dynamic games with coupling chance constraints, where agents have incomplete information about uncertainties satisfying a concentration of measure property. This problem, in general, is…

Systems and Control · Electrical Eng. & Systems 2026-02-06 Seyed Shahram Yadollahi , Hamed Kebriaei , Sadegh Soudjani

We consider in discrete time, a general class of sequential stochastic dynamic games with asymmetric information with the following features. The underlying system has Markovian dynamics controlled by the agents' joint actions. Each agent's…

Multiagent Systems · Computer Science 2023-01-16 Yi Ouyang , Hamidreza Tavafoghi , Demosthenis Teneketzis

We study a subclass of $n$-player stochastic games, namely, stochastic games with independent chains and unknown transition matrices. In this class of games, players control their own internal Markov chains whose transitions do not depend…

Computer Science and Game Theory · Computer Science 2023-12-05 Tiancheng Qin , S. Rasoul Etesami

This paper considers information sharing in a multi-player repeated game. Every round, each player observes a subset of components of a random vector and then takes a control action. The utility earned by each player depends on the full…

Optimization and Control · Mathematics 2014-12-31 Michael J. Neely

Discrete-time replicator map is a prototype of evolutionary selection game dynamical models that have been very successful across disciplines in rendering insights into the attainment of the equilibrium outcomes, like the Nash equilibrium…

Populations and Evolution · Quantitative Biology 2021-02-22 Archan Mukhopadhyay , Sagar Chakraborty

No-regret self-play learning dynamics have become one of the premier ways to solve large-scale games in practice. Accelerating their convergence via improving the regret of the players over the naive $O(\sqrt{T})$ bound after $T$ rounds has…

Machine Learning · Computer Science 2025-02-26 Shinji Ito , Haipeng Luo , Taira Tsuchiya , Yue Wu

Game theory has emerged as a powerful framework for modeling a large range of multi-agent scenarios. Many algorithmic solutions require discrete, finite games with payoffs that have a closed-form specification. In contrast, many real-world…

Computer Science and Game Theory · Computer Science 2018-06-13 Abdullah Al-Dujaili , Erik Hemberg , Una-May O'Reilly

A long-standing open problem in algorithmic game theory asks whether or not there is a polynomial time algorithm to compute a Nash equilibrium in a random bimatrix game. We study random win-lose games, where the entries of the $n\times n$…

Computer Science and Game Theory · Computer Science 2025-10-16 Andrea Collevecchio , Gabor Lugosi , Adrian Vetta , Rui-Ray Zhang

The design and performance analysis of bandit algorithms in the presence of stage-wise safety or reliability constraints has recently garnered significant interest. In this work, we consider the linear stochastic bandit problem under…

Machine Learning · Computer Science 2020-03-03 Ahmadreza Moradipari , Sanae Amani , Mahnoosh Alizadeh , Christos Thrampoulidis

In stochastic Nash equilibrium problems (SNEPs), it is natural for players to be uncertain about their complex environments and have multi-dimensional unknown parameters in their models. Among various SNEPs, this paper focuses on locally…

Optimization and Control · Mathematics 2022-04-06 Yuanhanqing Huang , Jianghai Hu

Thompson Sampling is one of the oldest heuristics for multi-armed bandit problems. It is a randomized algorithm based on Bayesian ideas, and has recently generated significant interest after several studies demonstrated it to have better…

Machine Learning · Computer Science 2012-09-18 Shipra Agrawal , Navin Goyal

In this paper, we consider a learning problem among non-cooperative agents interacting in a time-varying system. Specifically, we focus on repeated linear quadratic network games, in which the network of interactions changes with time and…

Computer Science and Game Theory · Computer Science 2023-10-23 Feras Al Taha , Kiran Rokade , Francesca Parise

We study the performance of the Thompson Sampling algorithm for logistic bandit problems. In this setting, an agent receives binary rewards with probabilities determined by a logistic function, $\exp(\beta \langle a, \theta…

Machine Learning · Statistics 2025-02-21 Amaury Gouverneur , Borja Rodríguez-Gálvez , Tobias J. Oechtering , Mikael Skoglund

This study investigates differential games with motion-payoff uncertainty in continuous-time settings. We propose a framework where players update their beliefs about uncertain parameters using continuous Bayesian updating. Theoretical…

Multiagent Systems · Computer Science 2025-09-16 Jiangjing Zhou , Ovanes Petrosian , Ye Zhang , Hongwei Gao

We consider a subclass of $n$-player stochastic games, in which players have their own internal state/action spaces while they are coupled through their payoff functions. It is assumed that players' internal chains are driven by independent…

Machine Learning · Computer Science 2023-03-23 S. Rasoul Etesami

This paper studies the stochastic linear bandit problem, where a decision-maker chooses actions from possibly time-dependent sets of vectors in $\mathbb{R}^d$ and receives noisy rewards. The objective is to minimize regret, the difference…

Machine Learning · Computer Science 2023-04-24 Nima Hamidi , Mohsen Bayati

We consider the problem of finding stationary Nash equilibria (NE) in a finite discounted general-sum stochastic game. We first generalize a non-linear optimization problem from Filar and Vrieze [2004] to a $N$-player setting and break down…

Computer Science and Game Theory · Computer Science 2015-07-06 H. L Prasad , L. A. Prashanth , Shalabh Bhatnagar

This paper considers repeated games in which one player has more information about the game than the other players. In particular, we investigate repeated two-player zero-sum games where only the column player knows the payoff matrix A of…

Computer Science and Game Theory · Computer Science 2023-02-16 Le Cong Dinh , Long Tran-Thanh , Tri-Dung Nguyen , Alain B. Zemkoho

We study discrete-time mean-field Markov games with infinite numbers of agents where each agent aims to minimize its ergodic cost. We consider the setting where the agents have identical linear state transitions and quadratic cost…

Optimization and Control · Mathematics 2019-10-17 Zuyue Fu , Zhuoran Yang , Yongxin Chen , Zhaoran Wang