English
Related papers

Related papers: A Payoff-Based Policy Gradient Method in Stochasti…

200 papers

There exist a number of reinforcement learning algorithms which learnby climbing the gradient of expected reward. Their long-runconvergence has been proved, even in partially observableenvironments with non-deterministic actions, and…

Machine Learning · Computer Science 2013-01-14 Lex Weaver , Nigel Tao

We examine the use of different randomisation policies for stochastic gradient algorithms used in sampling, based on first-order (or overdamped) Langevin dynamics, the most popular of which is known as Stochastic Gradient Langevin Dynamics.…

Numerical Analysis · Mathematics 2025-12-16 Luke Shaw , Peter A. Whalley

We consider a general-sum N-player linear-quadratic game with stochastic dynamics over a finite horizon and prove the global convergence of the natural policy gradient method to the Nash equilibrium. In order to prove the convergence of the…

Optimization and Control · Mathematics 2022-08-16 Ben Hambly , Renyuan Xu , Huining Yang

This paper proposes a payoff perturbation technique for the Mirror Descent (MD) algorithm in games where the gradient of the payoff functions is monotone in the strategy profile space, potentially containing additive noise. The optimistic…

Computer Science and Game Theory · Computer Science 2024-06-25 Kenshi Abe , Kaito Ariu , Mitsuki Sakamoto , Atsushi Iwasaki

We consider a distributed stochastic approximation (SA) scheme for computing an equilibrium of a stochastic Nash game. Standard SA schemes employ diminishing steplength sequences that are square summable but not summable. Such requirements…

Optimization and Control · Mathematics 2013-03-20 Farzad Yousefian , Angelia Nedich , Uday V. Shanbhag

In this paper, we propose a passivity-based methodology for analysis and design of reinforcement learning in multi-agent finite games. Starting from a known exponentially-discounted reinforcement learning scheme, we show that convergence to…

Optimization and Control · Mathematics 2024-10-30 Bolin Gao , Lacra Pavel

We propose a simple, general and effective technique, Reward Randomization for discovering diverse strategic policies in complex multi-agent games. Combining reward randomization and policy gradient, we derive a new algorithm,…

Artificial Intelligence · Computer Science 2021-03-15 Zhenggang Tang , Chao Yu , Boyuan Chen , Huazhe Xu , Xiaolong Wang , Fei Fang , Simon Du , Yu Wang , Yi Wu

We study a setting in which two players play a (possibly approximate) Nash equilibrium of a bimatrix game, while a learner observes only their actions and has no knowledge of the equilibrium or the underlying game. A natural question is…

Computer Science and Game Theory · Computer Science 2026-05-27 Annalisa Barbara , Riccardo Poiani , Martino Bernasconi , Andrea Celli

Policy gradient methods, which have been extensively studied in the last decade, offer an effective and efficient framework for reinforcement learning problems. However, their performances can often be unsatisfactory, suffering from…

Machine Learning · Computer Science 2026-01-27 Shihab Ahmed , El Houcine Bergou , Aritra Dutta , Yue Wang

In this paper,we consider the restless bandit problem, which is one of the most well-studied generalizations of the celebrated stochastic multi-armed bandit problem in decision theory. However, it is known be PSPACE-Hard to approximate to…

Machine Learning · Computer Science 2011-04-29 Quan Liu , Kehao Wang , Lin Chen

Learning or estimating game models from data typically entails inducing separate models for each setting, even if the games are parametrically related. In empirical mechanism design, for example, this approach requires learning a new game…

Computer Science and Game Theory · Computer Science 2026-05-05 Madelyn Gatchel , Michael P. Wellman

Many high-stakes decision-making problems, such as those found within cybersecurity and economics, can be modeled as competitive resource allocation games. In these games, multiple players must allocate limited resources to overcome their…

Computer Science and Game Theory · Computer Science 2024-01-10 N'yoma Diamond , Fabricio Murai

We investigate the resolution of second-order, potential, and monotone mean field games with the generalized conditional gradient algorithm, an extension of the Frank-Wolfe algorithm. We show that the method is equivalent to the fictitious…

Optimization and Control · Mathematics 2023-08-22 Pierre Lavigne , Laurent Pfeiffer

We introduce the Colonel Blotto game with favoritism, an extension of the famous Colonel Blotto game where the winner-determination rule is generalized to include pre-allocations and asymmetry of the players' resources effectiveness on each…

Computer Science and Game Theory · Computer Science 2021-06-02 Dong Quan Vu , Patrick Loiseau

This paper studies the last-iterate convergence properties of the exponential weights algorithm with constant learning rates. We consider a repeated interaction in discrete time, where each player uses an exponential weights algorithm…

Artificial Intelligence · Computer Science 2024-07-10 Maurizio d'Andrea , Fabien Gensbittel , Jérôme Renault

We study a subclass of $n$-player stochastic games, namely, stochastic games with independent chains and unknown transition matrices. In this class of games, players control their own internal Markov chains whose transitions do not depend…

Computer Science and Game Theory · Computer Science 2023-12-05 Tiancheng Qin , S. Rasoul Etesami

We introduce and study incentive equilibria for multi-player meanpayoff games. Incentive equilibria generalise well-studied solution concepts such as Nash equilibria and leader equilibria (also known as Stackelberg equilibria). Recall that…

Computer Science and Game Theory · Computer Science 2015-11-03 Anshul Gupta , M. S. Krishna Deepak , Bharath Kumar Padarthi , Sven Schewe , Ashutosh Trivedi

Zero-sum stochastic games are easy to solve as they can be cast as simple Markov decision processes. This is however not the case with general-sum stochastic games. A fairly general optimization problem formulation is available for…

Machine Learning · Computer Science 2015-07-02 H. L. Prasad , Shalabh Bhatnagar

We consider the computation of an equilibrium of a stochastic Nash equilibrium problem, where the player objectives are assumed to be $L_0$-Lipschitz continuous and convex given rival decisions with convex and closed player-specific…

Optimization and Control · Mathematics 2025-10-29 Luke Marrinan , Farzad Yousefian , Uday V. Shanbhag

Real world applications such as economics and policy making often involve solving multi-agent games with two unique features: (1) The agents are inherently asymmetric and partitioned into leaders and followers; (2) The agents have different…

Machine Learning · Computer Science 2021-11-04 Yu Bai , Chi Jin , Huan Wang , Caiming Xiong