English
Related papers

Related papers: Two-Player Zero-Sum Games with Bandit Feedback

200 papers

We focus on the design of algorithms for finding equilibria in 2-player zero-sum games. Although it is well known that such problems can be solved by a single linear program, there has been a surge of interest in recent years for simpler…

Computer Science and Game Theory · Computer Science 2025-02-03 Michail Fasoulakis , Evangelos Markakis , Giorgos Roussakis , Christodoulos Santorinaios

We consider a number of questions related to tradeoffs between reward and regret in repeated gameplay between two agents. To facilitate this, we introduce a notion of $\textit{generalized equilibrium}$ which allows for asymmetric regret…

Computer Science and Game Theory · Computer Science 2023-12-19 William Brown , Jon Schneider , Kiran Vodrahalli

Online learning algorithms that minimize regret provide strong guarantees in situations that involve repeatedly making decisions in an uncertain environment, e.g. a driver deciding what route to drive to work every day. While regret…

Computer Science and Game Theory · Computer Science 2013-09-06 Jeremiah Blocki , Nicolas Christin , Anupam Datta , Arunesh Sinha

We suggest a general method for inferring players' values from their actions in repeated games. The method extends and improves upon the recent suggestion of (Nekipelov et al., EC 2015) and is based on the assumption that players are more…

Computer Science and Game Theory · Computer Science 2017-02-17 Noam Nisan , Gali Noti

Partial monitoring games are repeated games where the learner receives feedback that might be different from adversary's move or even the reward gained by the learner. Recently, a general model of combinatorial partial monitoring (CPM)…

Computer Science and Game Theory · Computer Science 2016-08-24 Sougata Chaudhuri , Ambuj Tewari

Two-team zero-sum games are one of the most important paradigms in game theory. In this paper, we focus on finding an unexploitable equilibrium in large team games. An unexploitable equilibrium is a worst-case policy, where members in the…

Computer Science and Game Theory · Computer Science 2024-03-04 Naming Liu , Mingzhi Wang , Youzhi Zhang , Yaodong Yang , Bo An , Ying Wen

We propose a novel algorithm for multi-player multi-armed bandits without collision sensing information. Our algorithm circumvents two problems shared by all state-of-the-art algorithms: it does not need as an input a lower bound on the…

Machine Learning · Statistics 2022-06-07 Wei Huang , Richard Combes , Cindy Trinh

There have been extensive studies on learning in zero-sum games, focusing on the analysis of the existence and algorithmic convergence of Nash equilibrium (NE). Existing studies mainly focus on symmetric games where the strategy spaces of…

Computer Science and Game Theory · Computer Science 2025-02-11 Yuheng Li , Panpan Wang , Haipeng Chen

We study online optimization methods for zero-sum games, a fundamental problem in adversarial learning in machine learning, economics, and many other domains. Traditional methods approximate Nash equilibria (NE) using either regret-based…

Computer Science and Game Theory · Computer Science 2025-07-16 Taemin Kim , James P. Bailey

This paper proposes a game-theoretic approach to address the problem of optimal sensor placement against an adversary in uncertain networked control systems. The problem is formulated as a zero-sum game with two players, namely a malicious…

Systems and Control · Electrical Eng. & Systems 2023-01-13 Anh Tung Nguyen , Sribalaji C. Anand , André M. H. Teixeira

We study online reinforcement learning in linear Markov decision processes with adversarial losses and bandit feedback, without prior knowledge on transitions or access to simulators. We introduce two algorithms that achieve improved regret…

Machine Learning · Computer Science 2023-10-19 Haolin Liu , Chen-Yu Wei , Julian Zimmert

This paper considers an online multi-player resource-sharing game with bandit feedback. Multiple players choose from a finite collection of resources in a time slotted system. In each time slot, each resource brings a random reward that is…

Computer Science and Game Theory · Computer Science 2025-02-18 Mevan Wijewardena , Michael. J Neely

Nash equilibria provide a principled framework for modeling interactions in multi-agent decision-making and control. However, many equilibrium-seeking methods implicitly assume that each agent has access to the other agents' objectives and…

Computer Science and Game Theory · Computer Science 2026-03-19 Mahdis Rabbani , Navid Mojahed , Shima Nazari

We consider the non-stochastic version of the (cooperative) multi-player multi-armed bandit problem. The model assumes no communication at all between the players, and furthermore when two (or more) players select the same action this…

Machine Learning · Computer Science 2019-05-03 Sébastien Bubeck , Yuanzhi Li , Yuval Peres , Mark Sellke

We study stochastic linear optimization problem with bandit feedback. The set of arms take values in an $N$-dimensional space and belong to a bounded polyhedron described by finitely many linear inequalities. We provide a lower bound for…

Machine Learning · Computer Science 2015-09-29 Manjesh K. Hanawal , Amir Leshem , Venkatesh Saligrama

In this research, we investigate the high-dimensional linear contextual bandit problem where the number of features $p$ is greater than the budget $T$, or it may even be infinite. Differing from the majority of previous works in this field,…

Machine Learning · Statistics 2025-06-27 Junpei Komiyama , Masaaki Imaizumi

We consider two agents playing simultaneously the same stochastic three-armed bandit problem. The two agents are cooperating but they cannot communicate. We propose a strategy with no collisions at all between the players (with very high…

Computer Science and Game Theory · Computer Science 2020-07-13 Sébastien Bubeck , Thomas Budzinski

Classic no-regret multi-armed bandit algorithms, including the Upper Confidence Bound (UCB), Hedge, and EXP3, are inherently unfair by design. Their unfairness stems from their objective of playing the most rewarding arm as frequently as…

Machine Learning · Computer Science 2024-05-14 Abhishek Sinha

Research on the multi-armed bandit problem has studied the trade-off of exploration and exploitation in depth. However, there are numerous applications where the cardinal absolute-valued feedback model (e.g. ratings from one to five) is not…

Machine Learning · Computer Science 2018-12-12 Lennard Hilgendorf

Scale-invariance in games has recently emerged as a widely valued desirable property. Yet, almost all fast convergence guarantees in learning in games require prior knowledge of the utility scale. To address this, we develop learning…

Computer Science and Game Theory · Computer Science 2026-02-13 Taira Tsuchiya , Haipeng Luo , Shinji Ito
‹ Prev 1 4 5 6 7 8 10 Next ›