中文
相关论文

相关论文: Mean Field Equilibrium in Multi-Armed Bandit Game …

200 篇论文

Stochastic games provide a framework for interactions among multiple agents and enable a myriad of applications. In these games, agents decide on actions simultaneously, the state of every agent moves to the next state, and each agent…

机器学习 · 计算机科学 2019-10-10 Mridul Agarwal , Vaneet Aggarwal , Arnob Ghosh , Nilay Tiwari

The standard solution concept for stochastic games is Markov perfect equilibrium (MPE); however, its computation becomes intractable as the number of players increases. Instead, we consider mean field equilibrium (MFE) that has been…

理论经济学 · 经济学 2020-06-05 Bar Light , Gabriel Weintraub

We consider a multi-agent Markov strategic interaction over an infinite horizon where agents can be of multiple types. We model the strategic interaction as a mean-field game in the asymptotic limit when the number of agents of each type…

多智能体系统 · 计算机科学 2021-01-01 Arnob Ghosh , Vaneet Aggarwal

Mean Field Games (MFG) are the class of games with a very large number of agents and the standard equilibrium concept is a Mean Field Equilibrium (MFE). Algorithms for learning MFE in dynamic MFGs are unknown in general. Our focus is on an…

最优化与控制 · 数学 2021-02-02 Kiyeob Lee , Desik Rengarajan , Dileep Kalathil , Srinivas Shakkottai

This paper considers mean field games in a multi-agent Markov decision process (MDP) framework. Each player has a continuum state and binary action, and benefits from the improvement of the condition of the overall population. Based on an…

最优化与控制 · 数学 2021-01-05 Minyi Huang , Yan Ma

For a wireless avionics communication system, a Multi-arm bandit game is mathematically formulated, which includes channel states, strategies, and rewards. The simple case includes only two agents sharing the spectrum which is fully studied…

信号处理 · 电气工程与系统科学 2017-11-15 Jingyang Lu , Lun Li , Dan Shen , Genshe Chen , Bin Jia , Erik Blasch , Khanh Pham

The restless bandit problem is one of the most well-studied generalizations of the celebrated stochastic multi-armed bandit problem in decision theory. In its ultimate generality, the restless bandit problem is known to be PSPACE-Hard to…

数据结构与算法 · 计算机科学 2009-02-03 Sudipto Guha , Kamesh Munagala , Peng Shi

Multi-agent reinforcement learning methods have shown remarkable potential in solving complex multi-agent problems but mostly lack theoretical guarantees. Recently, mean field control and mean field games have been established as a…

机器学习 · 计算机科学 2021-12-20 Kai Cui , Anam Tahir , Mark Sinzger , Heinz Koeppl

The stochastic multi-armed bandit (MAB) problem is a common model for sequential decision problems. In the standard setup, a decision maker has to choose at every instant between several competing arms, each of them provides a scalar random…

机器学习 · 统计学 2021-10-27 Asaf Cassel , Shie Mannor , Assaf Zeevi

The multi-armed bandit (MAB) model is one of the most classical models to study decision-making in an uncertain environment. In this model, a player chooses one of $K$ possible arms of a bandit machine to play at each time step, where the…

机器学习 · 计算机科学 2023-06-13 Bo Li , Chi Ho Yeung

We consider learning approximate Nash equilibria for discrete-time mean-field games with nonlinear stochastic state dynamics subject to both average and discounted costs. To this end, we introduce a mean-field equilibrium (MFE) operator,…

系统与控制 · 电气工程与系统科学 2022-11-11 Berkay Anahtarcı , Can Deha Karıksız , Naci Saldi

This paper considers mean field games in a multi-agent Markov decision process (MDP) framework. Each player has a continuum state and binary action. By active control, a player can bring its state to a resetting point. All players are…

最优化与控制 · 数学 2017-01-25 Minyi Huang , Yan Ma

We study a class of stochastic dynamic games that exhibit strategic complementarities between players; formally, in the games we consider, the payoff of a player has increasing differences between her own state and the empirical…

计算机科学与博弈论 · 计算机科学 2010-12-13 Sachin Adlakha , Ramesh Johari

The multi-armed bandit (MAB) problem is a classic example of the exploration-exploitation dilemma. It is concerned with maximising the total rewards for a gambler by sequentially pulling an arm from a multi-armed slot machine where each arm…

机器学习 · 统计学 2018-05-16 Xue Lu , Niall Adams , Nikolas Kantas

Motivated by a number of real-world applications from domains like healthcare and sustainable transportation, in this paper we study a scenario of repeated principal-agent games within a multi-armed bandit (MAB) framework, where: the…

机器学习 · 计算机科学 2023-05-09 Ilgin Dogan , Zuo-Jun Max Shen , Anil Aswani

We consider a class of continuous-time dynamic games involving a large number of players. Each player selects actions from a finite set and evolves through a finite set of states. State transitions occur stochastically and depend on the…

系统与控制 · 电气工程与系统科学 2025-11-12 Leonardo Pedroso , Andrea Agazzi , W. P. M. H. Heemels , Mauro Salazar

We introduce a mean field model for optimal holding of a representative agent of her peers as a natural expected scaling limit from the corresponding $N-$agent model. The induced mean field dynamics appear naturally in a form which is not…

最优化与控制 · 数学 2022-04-05 Mao Fabrice Djete , Nizar Touzi

Mean field games formalize dynamic games with a continuum of players and explicit interaction where the players can have heterogeneous states. As they additionally yield approximate equilibria of corresponding $N$-player games, they are of…

最优化与控制 · 数学 2020-01-09 Berenice Anne Neumann

In this paper, we consider a finite horizon, non-stationary, mean field games (MFG) with a large population of homogeneous players, sequentially making strategic decisions, where each player is affected by other players through an aggregate…

系统与控制 · 电气工程与系统科学 2020-04-07 Rajesh K Mishra , Deepanshu Vasal , Sriram Vishwanath

We study the multi-armed bandit (MAB) problem with composite and anonymous feedback. In this model, the reward of pulling an arm spreads over a period of time (we call this period as reward interval) and the player receives partial rewards…

机器学习 · 计算机科学 2020-12-16 Siwei Wang , Haoyun Wang , Longbo Huang
‹ 上一页 1 2 3 10 下一页 ›