中文
相关论文

相关论文: Provable Sample Complexity Guarantees for Learning…

200 篇论文

In this paper we investigate the Follow the Regularized Leader dynamics in sequential imperfect information games (IIG). We generalize existing results of Poincar\'e recurrence from normal-form games to zero-sum two-player imperfect…

A classic model to study strategic decision making in multi-agent systems is the normal-form game. This model can be generalised to allow for an infinite number of pure strategies leading to continuous games. Multi-objective normal-form…

计算机科学与博弈论 · 计算机科学 2023-03-02 Willem Röpke , Carla Groenland , Roxana Rădulescu , Ann Nowé , Diederik M. Roijers

$ $This paper addresses the inverse problem for Linear-Quadratic (LQ) nonzero-sum $N$-player differential games, where the goal is to learn parameters of an unknown cost function for the game, called observed, given the demonstrated…

最优化与控制 · 数学 2024-10-28 Emin Martirosyan , Ming Cao

We present a computational formulation for the approximate version of several variational inequality problems, investigating their computational complexity and establishing PPAD-completeness. Examining applications in computational game…

计算复杂性 · 计算机科学 2024-11-08 Bruce M. Kapron , Koosha Samieefar

Reinforcement Learning Algorithms (RLA) are useful machine learning tools to understand how decision makers react to signals. It is known that RLA converge towards the pure Nash Equilibria (NE) of finite congestion games and more generally,…

计算机科学与博弈论 · 计算机科学 2021-11-15 Benoît Sohet , Yezekael Hayel , Olivier Beaude , Alban Jeandin

The Kelly or proportional allocation mechanism is a simple and efficient auction-based scheme that distributes an infinitely divisible resource proportionally to the agents bids. When agents are aware of the allocation rule, their…

计算机科学与博弈论 · 计算机科学 2026-03-27 Younes Ben Mazziane , Cleque-Marlain Mboulou Moutoubi , Eitan Altman , Francesco De Pellegrini

This paper considers the design of non-truthful mechanisms from samples. We identify a parameterized family of mechanisms with strategically simple winner-pays-bid, all-pay, and truthful payment formats. In general (not necessarily…

计算机科学与博弈论 · 计算机科学 2019-06-26 Jason Hartline , Samuel Taggart

In this paper, we examine the convergence landscape of multi-agent learning under uncertainty. Specifically, we analyze two stochastic models of regularized learning in continuous games -- one in continuous and one in discrete time with the…

计算机科学与博弈论 · 计算机科学 2025-12-10 Kyriakos Lotidis , Panayotis Mertikopoulos , Nicholas Bambos , Jose Blanchet

A growing line of work reframes preference-based fine-tuning of large language models game-theoretically: Nash Learning from Human Feedback (NLHF) recasts the problem as a zero-sum game over policies. However, optimization is over expected…

计算机科学与博弈论 · 计算机科学 2026-05-14 Max Horwitz , Jake Gonzales , Eric Mazumdar , Lillian J. Ratliff

Continuous-time empirical dynamic discrete choice games offer notable computational advantages over discrete-time models. This paper addresses remaining computational and econometric challenges to further improve both model solution and…

计量经济学 · 经济学 2025-11-11 Jason R. Blevins

This paper addresses policy learning in non-stationary environments and games with continuous actions. Rather than the classical reward maximization mechanism, inspired by the ideas of follow-the-regularized-leader (FTRL) and mirror descent…

机器学习 · 计算机科学 2022-08-22 Rong-Jun Qin , Fan-Ming Luo , Hong Qian , Yang Yu

We investigate uniformity properties of strategies. These properties involve sets of plays in order to express useful constraints on strategies that are not \mu-calculus definable. Typically, we can state that a strategy is…

计算机科学与博弈论 · 计算机科学 2013-03-05 Bastien Maubert , Sophie Pinchinat , Laura Bozzelli

We introduce a framework for stochastic games on large sparse graphs, covering continuous-time and discrete-time dynamic games as well as static games. Players are indexed by the vertices of simple, locally finite graphs, allowing both…

最优化与控制 · 数学 2026-02-27 Eyal Neuman , Sturmius Tuschmann

Consider a 2-player normal-form game repeated over time. We introduce an adaptive learning procedure, where the players only observe their own realized payoff at each stage. We assume that agents do not know their own payoff function, and…

计算机科学与博弈论 · 计算机科学 2013-06-13 Mario Bravo , Mathieu Faure

Multi-agent reinforcement learning has made substantial empirical progresses in solving games with a large number of players. However, theoretically, the best known sample complexity for finding a Nash equilibrium in general-sum games…

机器学习 · 计算机科学 2022-04-01 Ziang Song , Song Mei , Yu Bai

We study best-response type learning dynamics for zero-sum polymatrix games under two information settings. The two settings are distinguished by the type of information that each player has about the game and their opponents' strategy. The…

最优化与控制 · 数学 2025-08-13 Fathima Zarin Faizal , Asuman Ozdaglar , Martin J. Wainwright

We present an algorithm that computes approximate pure Nash equilibria in a broad class of constraint satisfaction games that generalize the well-known cut and party affiliation games. Our results improve previous ones by Bhalgat et al.~(EC…

计算机科学与博弈论 · 计算机科学 2014-02-17 Ioannis Caragiannis , Angelo Fanelli , Nick Gravin

We study the problem of learning in zero-sum matrix games with repeated play and bandit feedback. Specifically, we focus on developing uncoupled algorithms that guarantee, without communication between players, the convergence of the…

机器学习 · 计算机科学 2026-04-20 Côme Fiegel , Pierre Ménard , Tadashi Kozuno , Michal Valko , Vianney Perchet

Motivated by the scarcity of accurate payoff feedback in practical applications of game theory, we examine a class of learning dynamics where players adjust their choices based on past payoff observations that are subject to noise and…

最优化与控制 · 数学 2016-06-03 Mario Bravo , Panayotis Mertikopoulos

We study the long-term behavior of the fictitious play process in repeated extensive-form games of imperfect information with perfect recall. Each player maintains incorrect beliefs that the moves at all information sets, except the one at…

计算机科学与博弈论 · 计算机科学 2025-04-28 Jason Castiglione , Gürdal Arslan