中文
相关论文

相关论文: Two families of indexable partially observable res…

200 篇论文

We consider the problem of learning the optimal threshold policy for control problems. Threshold policies make control decisions by evaluating whether an element of the system state exceeds a certain threshold, whose value is determined by…

机器学习 · 计算机科学 2022-09-30 Khaled Nakhleh , I-Hong Hou

In restless bandits, a central agent is tasked with optimally distributing limited resources across several bandits (arms), with each arm being a Markov decision process. In this work, we generalize the traditional restless bandits problem…

机器学习 · 计算机科学 2026-02-20 Nima Akbarzadeh , Yossiri Adulyasak , Erick Delage

This paper studies restless multi-armed bandit (RMAB) problems with unknown arm transition dynamics but with known correlated arm features. The goal is to learn a model to predict transition dynamics given features, where the Whittle index…

机器学习 · 计算机科学 2023-08-15 Kai Wang , Shresth Verma , Aditya Mate , Sanket Shah , Aparna Taneja , Neha Madhiwalla , Aparna Hegde , Milind Tambe

In this paper, we consider a queueing system with multiple channels (or servers) and multiple classes of users. We aim at allocating the available channels among the users in such a way to minimize the expected total average queue length of…

网络与互联网体系结构 · 计算机科学 2019-02-07 Saad Kriouile , Maialen Larranaga , Mohamad Assaad

This paper studies a class of constrained restless multi-armed bandits (CRMAB). The constraints are in the form of time varying set of actions (set of available arms). This variation can be either stochastic or semi-deterministic. Given a…

系统与控制 · 计算机科学 2021-09-07 Kesav Kaza , Rahul Meshram , Varun Mehta , S. N. Merchant

The optimal scheduling problem in single-server queueing systems is a classic problem in queueing theory. The Gittins index policy is known to be the optimal preemptive nonanticipating policy (both for the open version of the problem with…

性能 · 计算机科学 2021-03-22 Samuli Aalto

We evaluate the performance of Whittle index policy for restless Markovian bandits, when the number of bandits grows. It is proven in [30] that this performance is asymptotically optimal if the bandits are indexable and the associated…

性能 · 计算机科学 2020-12-17 Nicolas Gast , Bruno Gaujal , Chen Yan

We study a resource allocation problem with varying requests, and with resources of limited capacity shared by multiple requests. It is modeled as a set of heterogeneous Restless Multi-Armed Bandit Problems (RMABPs) connected by constraints…

最优化与控制 · 数学 2020-03-30 Jing Fu , Bill Moran , Peter G. Taylor

This paper develops a polyhedral approach to the design, analysis, and computation of dynamic allocation indices for scheduling binary-action (engage/rest) Markovian stochastic projects which can change state when rested (restless bandits…

最优化与控制 · 数学 2023-04-05 José Niño-Mora

We consider finite-horizon restless bandits with multiple pulls per period, which play an important role in recommender systems, active learning, revenue management, and many other areas. While an optimal policy can be computed, in…

最优化与控制 · 数学 2021-07-27 Xiangyu Zhang , Peter I. Frazier

We study the Whittle index learning algorithm for restless multi-armed bandits (RMAB). We first present Q-learning algorithm and its variants -- speedy Q-learning (SQL), generalized speedy Q-learning (GSQL) and phase Q-learning (PhaseQL).…

机器学习 · 计算机科学 2024-09-11 Parvish Kakarapalli , Devendra Kayande , Rahul Meshram

In this paper,we consider the restless bandit problem, which is one of the most well-studied generalizations of the celebrated stochastic multi-armed bandit problem in decision theory. However, it is known be PSPACE-Hard to approximate to…

机器学习 · 计算机科学 2011-04-29 Quan Liu , Kehao Wang , Lin Chen

Restless Multi-Armed Bandits (RMABs) offer a powerful framework for solving resource constrained maximization problems. However, the formulation can be inappropriate for settings where the limiting constraint is a reward threshold rather…

数据结构与算法 · 计算机科学 2024-09-06 R. Teal Witter , Lisa Hellerstein

Restless and collapsing bandits are often used to model budget-constrained resource allocation in settings where arms have action-dependent transition probabilities, such as the allocation of health interventions among patients. However,…

机器学习 · 计算机科学 2023-07-20 Christine Herlihy , Aviva Prins , Aravind Srinivasan , John P. Dickerson

Restless Multi-Armed Bandits (RMABs) are powerful models for decision-making under uncertainty, yet classical formulations typically assume fixed dynamics, an assumption often violated in nonstationary environments. We introduce MARBLE…

机器学习 · 计算机科学 2026-04-13 Mohsen Amiri , Konstantin Avrachenkov , Ibtihal El Mimouni , Sindri Magnússon

This paper investigates the Restless Multi-Armed Bandit (RMAB) framework under individual penalty constraints to address resource allocation challenges in dynamic wireless networked environments. Unlike conventional RMAB models, our model…

机器学习 · 计算机科学 2026-04-20 Nida Zamir , I-Hong Hou

We study a finite-horizon restless multi-armed bandit problem with multiple actions, dubbed R(MA)^2B. The state of each arm evolves according to a controlled Markov decision process (MDP), and the reward of pulling an arm depends on both…

机器学习 · 计算机科学 2022-03-25 Guojun Xiong , Jian Li , Rahul Singh

We study a problem of information gathering in a social network with dynamically available sources and time varying quality of information. We formulate this problem as a restless multi-armed bandit (RMAB). In this problem, information…

系统与控制 · 计算机科学 2018-01-22 Varun Mehta , Rahul Meshram , Kesav Kaza , S. N. Merchant

The restless multi-armed bandit (RMAB) framework is a popular model with applications across a wide variety of fields. However, its solution is hindered by the exponentially growing state space (with respect to the number of arms) and the…

机器学习 · 计算机科学 2025-08-05 Gongpu Chen , Soung Chang Liew , Deniz Gunduz

We propose Streaming Bandits, a Restless Multi Armed Bandit (RMAB) framework in which heterogeneous arms may arrive and leave the system after staying on for a finite lifetime. Streaming Bandits naturally capture the health intervention…

机器学习 · 计算机科学 2022-02-17 Aditya Mate , Arpita Biswas , Christoph Siebenbrunner , Susobhan Ghosh , Milind Tambe