English
Related papers

Related papers: Optimism Without Regularization: Constant Regret i…

200 papers

Quantile (and, more generally, KL) regret bounds, such as those achieved by NormalHedge (Chaudhuri, Freund, and Hsu 2009) and its variants, relax the goal of competing against the best individual expert to only competing against a majority…

Machine Learning · Statistics 2021-11-09 Jeffrey Negrea , Blair Bilodeau , Nicolò Campolongo , Francesco Orabona , Daniel M. Roy

Two-player zero-sum games of infinite duration and their quantitative versions are used in verification to model the interaction between a controller (Eve) and its environment (Adam). The question usually addressed is that of the existence…

Computer Science and Game Theory · Computer Science 2021-10-29 Paul Hunter , Guillermo A. Pérez , Jean-François Raskin

We introduce the first best-of-both-worlds algorithm for contextual combinatorial semi-bandits that simultaneously guarantees $\widetilde{\mathcal{O}}(\sqrt{T})$ regret in the adversarial regime and $\widetilde{\mathcal{O}}(\ln T)$ regret…

Machine Learning · Statistics 2026-03-27 Mengmeng Li , Philipp J. Schneider , Jelisaveta Aleksić , Daniel Kuhn

We study the problem of determining an effective exploration strategy in static and non-linear optimization problems, which depend on an unknown scalar parameter to be learned from online collected noisy data. An optimal trade-off between…

Optimization and Control · Mathematics 2024-09-13 Ying Wang , Mirko Pasquini , Kévin Colin , Håkan Hjalmarsson

An ideal strategy in zero-sum games should not only grant the player an average reward no less than the value of Nash equilibrium, but also exploit the (adaptive) opponents when they are suboptimal. While most existing works in Markov games…

Machine Learning · Computer Science 2022-06-15 Qinghua Liu , Yuanhao Wang , Chi Jin

Extensive-form games are a common model for multiagent interactions with imperfect information. In two-player zero-sum games, the typical solution concept is a Nash equilibrium over the unconstrained strategy set for each player. In many…

Computer Science and Game Theory · Computer Science 2019-02-07 Trevor Davis , Kevin Waugh , Michael Bowling

We study the problem of learning the utility functions of no-regret learning agents in a repeated normal-form game. Differing from most prior literature, we introduce a principal with the power to observe the agents playing the game, send…

Computer Science and Game Theory · Computer Science 2026-05-13 Brian Hu Zhang , Tao Lin , Yiling Chen , Tuomas Sandholm

No-regret learning has been widely used to compute a Nash equilibrium in two-person zero-sum games. However, there is still a lack of regret analysis for network stochastic zero-sum games, where players competing in two subnetworks only…

Optimization and Control · Mathematics 2022-05-31 Shijie Huang , Jinlong Lei , Yiguang Hong

Discounted-sum games provide a formal model for the study of reinforcement learning, where the agent is enticed to get rewards early since later rewards are discounted. When the agent interacts with the environment, she may regret her…

Computer Science and Game Theory · Computer Science 2018-11-20 Michaël Cadilhac , Guillermo A. Pérez , Marie van den Bogaard

We investigate a nonstochastic bandit setting in which the loss of an action is not immediately charged to the player, but rather spread over the subsequent rounds in an adversarial way. The instantaneous loss observed by the player at the…

Machine Learning · Computer Science 2022-09-27 Nicolò Cesa-Bianchi , Tommaso Cesari , Roberto Colomboni , Claudio Gentile , Yishay Mansour

In this paper, a new method is proposed to compute the rolling Nash equilibrium of the time-invariant nonlinear two-person zero-sum differential games. The idea is to discretize the time to transform a differential game into a sequential…

Systems and Control · Electrical Eng. & Systems 2020-11-13 Wei Liao , Xiaohui Wei , Jizhou Lai

We study an online forecasting setting in which, over $T$ rounds, $N$ strategic experts each report a forecast to a mechanism, the mechanism selects one forecast, and then the outcome is revealed. In any given round, each expert has a…

Machine Learning · Computer Science 2025-02-18 Junpei Komiyama , Nishant A. Mehta , Ali Mortazavi

In this paper we establish efficient and \emph{uncoupled} learning dynamics so that, when employed by all players in a general-sum multiplayer game, the \emph{swap regret} of each player after $T$ repetitions of the game is bounded by…

Computer Science and Game Theory · Computer Science 2022-10-07 Ioannis Anagnostides , Gabriele Farina , Christian Kroer , Chung-Wei Lee , Haipeng Luo , Tuomas Sandholm

Imperfect information games (IIG) are games in which each player only partially observes the current game state. We study how to learn $\epsilon$-optimal strategies in a zero-sum IIG through self-play with trajectory feedback. We give a…

Machine Learning · Statistics 2023-02-16 Côme Fiegel , Pierre Ménard , Tadashi Kozuno , Rémi Munos , Vianney Perchet , Michal Valko

Model-based reinforcement learning (MBRL) agents typically learn world models by minimizing predictive loss. However, powerful RL optimizers inevitably exploit minor model inaccuracies, leading to simulator exploitation and a reality gap…

Machine Learning · Computer Science 2026-05-29 Christoph Dann , Yishay Mansour , Mehryar Mohri

Characterizing the performance of no-regret dynamics in multi-player games is a foundational problem at the interface of online learning and game theory. Recent results have revealed that when all players adopt specific learning algorithms,…

Computer Science and Game Theory · Computer Science 2023-11-28 Ioannis Anagnostides , Alkis Kalavasis , Tuomas Sandholm , Manolis Zampetakis

Partial-monitoring games constitute a mathematical framework for sequential decision making problems with imperfect feedback: The learner repeatedly chooses an action, opponent responds with an outcome, and then the learner suffers a loss…

Computer Science and Game Theory · Computer Science 2011-10-13 András Antos , Gábor Bartók , Dávid Pál , Csaba Szepesvári

In contrast to the classic formulation of partial monitoring, linear partial monitoring can model infinite outcome spaces, while imposing a linear structure on both the losses and the observations. This setting can be viewed as a…

Machine Learning · Computer Science 2026-01-15 Federico Di Gennaro , Khaled Eldowa , Nicolò Cesa-Bianchi

Mechanism design has found considerable application to the construction of agent-interaction protocols. In the standard setting, the type (e.g., utility function) of an agent is not known by other agents, nor is it known by the mechanism…

Computer Science and Game Theory · Computer Science 2012-07-19 Nathanael Hyafil , Craig Boutilier

Fictitious play (FP) is a history-based strategy to choose actions in normal-form games, where players best-respond to the empirical frequency of their opponents' past actions. While it is well-established that FP converges to the set of…

Computer Science and Game Theory · Computer Science 2026-04-10 Jaehong Moon
‹ Prev 1 8 9 10 Next ›