English
Related papers

Related papers: Adaptive Scheduling: A Reinforcement Learning Whit…

200 papers

Restless multi-armed bandits (RMABs) extend multi-armed bandits to allow for stateful arms, where the state of each arm evolves restlessly with different transitions depending on whether that arm is pulled. Solving RMABs requires…

Machine Learning · Computer Science 2023-11-21 Kai Wang , Lily Xu , Aparna Taneja , Milind Tambe

We study the Whittle index learning algorithm for restless multi-armed bandits (RMAB). We first present Q-learning algorithm and its variants -- speedy Q-learning (SQL), generalized speedy Q-learning (GSQL) and phase Q-learning (PhaseQL).…

Machine Learning · Computer Science 2024-09-11 Parvish Kakarapalli , Devendra Kayande , Rahul Meshram

This paper studies restless multi-armed bandit (RMAB) problems with unknown arm transition dynamics but with known correlated arm features. The goal is to learn a model to predict transition dynamics given features, where the Whittle index…

Machine Learning · Computer Science 2023-08-15 Kai Wang , Shresth Verma , Aditya Mate , Sanket Shah , Aparna Taneja , Neha Madhiwalla , Aparna Hegde , Milind Tambe

Whittle index policy is a heuristic to the intractable restless multi-armed bandits (RMAB) problem. Although it is provably asymptotically optimal, finding Whittle indices remains difficult. In this paper, we present Neural-Q-Whittle, a…

Machine Learning · Computer Science 2023-10-04 Guojun Xiong , Jian Li

Energy demands from data centers have surged and stressed the grid in recent years. Electric grids require balancing supply and demand every second, motivating demand response (reduction) from large loads, including data centers. This can…

Computational Engineering, Finance, and Science · Computer Science 2026-05-20 Yifu Ding , Zixi Chen , Thomas Magnanti

The restless multi-armed bandit (RMAB) framework is a popular approach to solving resource allocation problems in networked systems. In this paper, we study optimal resource allocation in RMABs facing unknown and non-stationary dynamics.…

Machine Learning · Computer Science 2026-04-22 Md Kamran Chowdhury Shisher , Vishrant Tripathi , Mung Chiang , Christopher G. Brinton

We study the problem of planning restless multi-armed bandits (RMABs) with multiple actions. This is a popular model for multi-agent systems with applications like multi-channel communication, monitoring and machine maintenance tasks, and…

Multiagent Systems · Computer Science 2023-03-01 Abheek Ghosh , Dheeraj Nagaraj , Manish Jain , Milind Tambe

A novel reinforcement learning algorithm is introduced for multiarmed restless bandits with average reward, using the paradigms of Q-learning and Whittle index. Specifically, we leverage the structure of the Whittle index policy to reduce…

Machine Learning · Computer Science 2021-09-22 Konstantin E. Avrachenkov , Vivek S. Borkar

We consider a class of restless multi-armed bandit problems (RMBP) that arises in dynamic multichannel access, user/server scheduling, and optimal activation in multi-agent systems. For this class of RMBP, we establish the indexability and…

Information Theory · Computer Science 2008-11-13 Keqin Liu , Qing Zhao

This paper investigates the Restless Multi-Armed Bandit (RMAB) framework under individual penalty constraints to address resource allocation challenges in dynamic wireless networked environments. Unlike conventional RMAB models, our model…

Machine Learning · Computer Science 2026-04-20 Nida Zamir , I-Hong Hou

Multi-action restless multi-armed bandits (RMABs) are a powerful framework for constrained resource allocation in which $N$ independent processes are managed. However, previous work only study the offline setting where problem dynamics are…

Machine Learning · Computer Science 2021-06-24 Jackson A. Killian , Arpita Biswas , Sanket Shah , Milind Tambe

A smart target, also referred to as a reactive target, can take maneuvering motions to hinder radar tracking. We address beam scheduling for tracking multiple smart targets in phased array radar networks. We aim to mitigate the performance…

Systems and Control · Electrical Eng. & Systems 2023-12-14 Yuhang Hao , Zengfu Wang , José Niño-Mora , Jing Fu , Min Yang , Quan Pan

This paper studies a class of constrained restless multi-armed bandits (CRMAB). The constraints are in the form of time varying set of actions (set of available arms). This variation can be either stochastic or semi-deterministic. Given a…

Systems and Control · Computer Science 2021-09-07 Kesav Kaza , Rahul Meshram , Varun Mehta , S. N. Merchant

Dynamic spectrum access problem is an important problem that allows a wireless sub-network to use channels temporarily unoccupied by the parent network for minimizing the spectrum waste. Previous work has shown that the sequential channel…

Optimization and Control · Mathematics 2026-05-15 Keqin Liu , Qizhen Jia , Yiying Zhang , Zhi Ding

This paper addresses an important class of restless multi-armed bandit (RMAB) problems that finds broad application in operations research, stochastic optimization, and reinforcement learning. There are $N$ independent Markov processes that…

Optimization and Control · Mathematics 2025-04-18 Keqin Liu

We consider a quantum switch with a finite number of quantum memory registers that aims to serve multipartite entanglement requests among $N$ users. We propose scheduling policies that aim to optimize the average number of requests served…

Information Theory · Computer Science 2026-03-25 Subhankar Banerjee , Stavros Mitrolaris , Sennur Ulukus

We consider a large-scale cyber network with N components (e.g., paths, servers, subnets). Each component is either in a healthy state (0) or an abnormal state (1). Due to random intrusions, the state of each component transits from 0 to 1…

Systems and Control · Computer Science 2011-12-02 Keqin Liu , Qing Zhao

The restless multi-armed bandit (RMAB) framework is a popular model with applications across a wide variety of fields. However, its solution is hindered by the exponentially growing state space (with respect to the number of arms) and the…

Machine Learning · Computer Science 2025-08-05 Gongpu Chen , Soung Chang Liew , Deniz Gunduz

We study a finite-horizon restless multi-armed bandit problem with multiple actions, dubbed R(MA)^2B. The state of each arm evolves according to a controlled Markov decision process (MDP), and the reward of pulling an arm depends on both…

Machine Learning · Computer Science 2022-03-25 Guojun Xiong , Jian Li , Rahul Singh

In many public health settings, it is important for patients to adhere to health programs, such as taking medications and periodic health checks. Unfortunately, beneficiaries may gradually disengage from such programs, which is detrimental…

Machine Learning · Computer Science 2021-07-26 Arpita Biswas , Gaurav Aggarwal , Pradeep Varakantham , Milind Tambe
‹ Prev 1 2 3 10 Next ›