中文
相关论文

相关论文: An Asymptotically Optimal Index Policy for Finite-…

200 篇论文

We study the problem of planning restless multi-armed bandits (RMABs) with multiple actions. This is a popular model for multi-agent systems with applications like multi-channel communication, monitoring and machine maintenance tasks, and…

多智能体系统 · 计算机科学 2023-03-01 Abheek Ghosh , Dheeraj Nagaraj , Manish Jain , Milind Tambe

Restless multi-armed bandits (RMABs) have been highly successful in optimizing sequential resource allocation across many domains. However, in many practical settings with highly scarce resources, where each agent can only receive at most…

多智能体系统 · 计算机科学 2025-01-13 Guojun Xiong , Haichuan Wang , Yuqi Pan , Saptarshi Mandal , Sanket Shah , Niclas Boehmer , Milind Tambe

This paper addresses an important class of restless multi-armed bandit (RMAB) problems that finds broad application in operations research, stochastic optimization, and reinforcement learning. There are $N$ independent Markov processes that…

最优化与控制 · 数学 2025-04-18 Keqin Liu

We consider the infinite-horizon, average-reward restless bandit problem in discrete time. We propose a new class of policies that are designed to drive a progressively larger subset of arms toward the optimal distribution. We show that our…

机器学习 · 计算机科学 2026-03-31 Yige Hong , Qiaomin Xie , Yudong Chen , Weina Wang

We study a finite-horizon restless multi-armed bandit problem with multiple actions, dubbed R(MA)^2B. The state of each arm evolves according to a controlled Markov decision process (MDP), and the reward of pulling an arm depends on both…

机器学习 · 计算机科学 2022-03-25 Guojun Xiong , Jian Li , Rahul Singh

We study a problem of information gathering in a social network with dynamically available sources and time varying quality of information. We formulate this problem as a restless multi-armed bandit (RMAB). In this problem, information…

系统与控制 · 计算机科学 2018-01-22 Varun Mehta , Rahul Meshram , Kesav Kaza , S. N. Merchant

Restless bandits are an important class of problems with applications in recommender systems, active learning, revenue management and other areas. We consider infinite-horizon discounted restless bandits with many arms where a fixed…

机器学习 · 计算机科学 2022-03-31 Xiangyu Zhang , Peter I. Frazier

We consider the discrete time infinite horizon average reward restless markovian bandit (RMAB) problem. We propose a \emph{model predictive control} based non-stationary policy with a rolling computational horizon $\tau$. At each time-slot,…

最优化与控制 · 数学 2025-06-06 Nicolas Gast , Dheeraj Narasimha

Restless multi-armed bandits (RMABs) provide a scalable framework for sequential decision-making under uncertainty, but classical formulations assume binary actions and a single global budget. Real-world settings, such as healthcare, often…

机器学习 · 计算机科学 2025-10-28 Himadri S. Pandey , Kai Wang , Gian-Gabriel P. Garcia

Partially observable restless multi-armed bandits have found numerous applications including in recommendation systems, communication systems, public healthcare outreach systems, and in operations research. We study multi-action partially…

机器学习 · 计算机科学 2025-09-03 Rahul Meshram , Kesav Kaza

We consider a general infinite horizon Heterogeneous Restless multi-armed Bandit (RMAB). Heterogeneity is a fundamental problem for many real-world systems largely because it resists many concentration arguments. In this paper, we assume…

最优化与控制 · 数学 2025-11-12 Dheeraj Narasimha , Nicolas Gast

For the model of constrained multi-armed bandit, we show that by construction there exists an index-based deterministic asymptotically optimal algorithm. The optimality is achieved by the convergence of the probability of choosing an…

最优化与控制 · 数学 2020-07-30 Hyeong Soo Chang

For the stochastic multi-armed bandit (MAB) problem from a constrained model that generalizes the classical one, we show that an asymptotic optimality is achievable by a simple strategy extended from the $\epsilon_t$-greedy strategy. We…

最优化与控制 · 数学 2018-05-04 Hyeong Soo Chang

We consider finite state restless multi-armed bandit problem. The decision maker can act on M bandits out of N bandits in each time step. The play of arm (active arm) yields state dependent rewards based on action and when the arm is not…

机器学习 · 计算机科学 2023-05-02 Vishesh Mittal , Rahul Meshram , Deepak Dev , Surya Prakash

In the classic Bayesian restless multi-armed bandit (RMAB) problem, there are $N$ arms, with rewards on all arms evolving at each time as Markov chains with known parameters. A player seeks to activate $K \geq 1$ arms at each time in order…

最优化与控制 · 数学 2011-12-25 Wenhan Dai , Yi Gai , Bhaskar Krishnamachari , Qing Zhao

We study the Lagrangian Index Policy (LIP) for restless multi-armed bandits with long-run average reward. In particular, we compare the performance of LIP with the performance of the Whittle Index Policy (WIP), both heuristic policies known…

机器学习 · 计算机科学 2026-01-01 Konstantin Avrachenkov , Vivek S. Borkar , Pratik Shah

The restless multi-armed bandit (RMAB) framework is a popular model with applications across a wide variety of fields. However, its solution is hindered by the exponentially growing state space (with respect to the number of arms) and the…

机器学习 · 计算机科学 2025-08-05 Gongpu Chen , Soung Chang Liew , Deniz Gunduz

We adopt an optimal-control framework for addressing the undiscounted infinite-horizon discrete-time restless $N$-armed bandit problem. Unlike most studies that rely on constructing policies based on the relaxed single-armed Markov Decision…

最优化与控制 · 数学 2024-03-19 Chen YAN

This paper studies a class of constrained restless multi-armed bandits (CRMAB). The constraints are in the form of time varying set of actions (set of available arms). This variation can be either stochastic or semi-deterministic. Given a…

系统与控制 · 计算机科学 2021-09-07 Kesav Kaza , Rahul Meshram , Varun Mehta , S. N. Merchant

We propose minimum empirical divergence (MED) policy for the multiarmed bandit problem. We prove asymptotic optimality of the proposed policy for the case of finite support models. In our setting, Burnetas and Katehakis has already proposed…

统计理论 · 数学 2011-11-21 Junya Honda , Akimichi Takemura
‹ 上一页 1 2 3 10 下一页 ›