English
Related papers

Related papers: Adaptive Scheduling: A Reinforcement Learning Whit…

200 papers

Restless Multi-Armed Bandits (RMABs) offer a powerful framework for solving resource constrained maximization problems. However, the formulation can be inappropriate for settings where the limiting constraint is a reward threshold rather…

Data Structures and Algorithms · Computer Science 2024-09-06 R. Teal Witter , Lisa Hellerstein

Scheduling in multi-channel wireless communication system presents formidable challenges in effectively allocating resources. To address these challenges, we investigate a multi-resource restless matching bandit (MR-RMB) model for…

Machine Learning · Computer Science 2024-08-21 Nida Zamir , I-Hong Hou

The exploration-exploitation dilemma has been a central challenge in reinforcement learning (RL) with complex model classes. In this paper, we propose a new algorithm, Monotonic Q-Learning with Upper Confidence Bound (MQL-UCB) for RL with…

Machine Learning · Computer Science 2025-10-06 Heyang Zhao , Jiafan He , Quanquan Gu

We study a problem of information gathering in a social network with dynamically available sources and time varying quality of information. We formulate this problem as a restless multi-armed bandit (RMAB). In this problem, information…

Systems and Control · Computer Science 2018-01-22 Varun Mehta , Rahul Meshram , Kesav Kaza , S. N. Merchant

This study introduces ContextWIN, a novel architecture that extends the Neural Whittle Index Network (NeurWIN) model to address Restless Multi-Armed Bandit (RMAB) problems with a context-aware approach. By integrating a mixture of experts…

Machine Learning · Computer Science 2024-10-15 Zhanqiu Guo , Wayne Wang

The Whittle index for restless bandits (two-action semi-Markov decision processes) provides an intuitively appealing optimal policy for controlling a single generic project that can be active (engaged) or passive (rested) at each decision…

Optimization and Control · Mathematics 2026-01-22 José Niño-Mora

Motivated by applications such as machine repair, project monitoring, and anti-poaching patrol scheduling, we study intervention planning of stochastic processes under resource constraints. This planning problem has previously been modeled…

Artificial Intelligence · Computer Science 2026-02-26 Arpita Biswas , Jackson A. Killian , Paula Rodriguez Diaz , Susobhan Ghosh , Milind Tambe

The Whittle index policy is a heuristic that has shown remarkably good performance (with guaranteed asymptotic optimality) when applied to the class of problems known as Restless Multi-Armed Bandit Problems (RMABPs). In this paper we…

Artificial Intelligence · Computer Science 2024-06-05 Francisco Robledo Relaño , Vivek Borkar , Urtzi Ayesta , Konstantin Avrachenkov

We study a sampling and transmission scheduling problem for multi-source remote estimation, where a scheduler determines when to take samples from multiple continuous-time Gauss-Markov processes and send the samples over multiple channels…

Networking and Internet Architecture · Computer Science 2024-04-17 Tasmeen Zaman Ornee , Yin Sun

We consider the problem of content caching at the wireless edge to serve a set of end users via unreliable wireless channels so as to minimize the average latency experienced by end users due to the constrained wireless edge cache capacity.…

Networking and Internet Architecture · Computer Science 2023-02-24 Guojun Xiong , Shufan Wang , Jian Li , Rahul Singh

We study the Whittle index learning algorithm for restless multi-armed bandits. We consider index learning algorithm with Q-learning. We first present Q-learning algorithm with exploration policies -- epsilon-greedy, softmax,…

Machine Learning · Computer Science 2024-09-10 Vishesh Mittal , Rahul Meshram , Surya Prakash

We consider a wireless network in which a source node needs to transmit a large file to a destination node. The direct wireless link between the source and the destination is assumed to be blocked. Multiple candidate relays are available to…

Networking and Internet Architecture · Computer Science 2025-08-29 Mandar R. Nalavade , Ravindra S. Tomar , Gaurav S. Kasbekar

Restless multi-armed bandits (RMAB) have been widely used to model sequential decision making problems with constraints. The decision maker (DM) aims to maximize the expected total reward over an infinite horizon under an "instantaneous…

Machine Learning · Computer Science 2023-12-25 Shufan Wang , Guojun Xiong , Jian Li

We consider the client selection problem in wireless Federated Learning (FL), with the objective of reducing the total required time to achieve a certain level of learning accuracy. Since the server cannot observe the clients' dynamic…

Machine Learning · Computer Science 2025-09-22 Qiyue Li , Yingxin Liu , Hang Qi , Jieping Luo , Zhizhang Liu , Jingjin Wu

Restless Multi-Armed Bandits (RMABs) are powerful models for decision-making under uncertainty, yet classical formulations typically assume fixed dynamics, an assumption often violated in nonstationary environments. We introduce MARBLE…

Machine Learning · Computer Science 2026-04-13 Mohsen Amiri , Konstantin Avrachenkov , Ibtihal El Mimouni , Sindri Magnússon

We propose Streaming Bandits, a Restless Multi Armed Bandit (RMAB) framework in which heterogeneous arms may arrive and leave the system after staying on for a finite lifetime. Streaming Bandits naturally capture the health intervention…

Machine Learning · Computer Science 2022-02-17 Aditya Mate , Arpita Biswas , Christoph Siebenbrunner , Susobhan Ghosh , Milind Tambe

The adoption of dynamic, self-learning solutions for real-time wireless network optimization has recently gained significant attention due to the limited adaptability of existing protocols. This paper investigates multi-armed bandit (MAB)…

Networking and Internet Architecture · Computer Science 2025-12-01 Miguel Casasnovas , Francesc Wilhelmi , Richard Combes , Maksymilian Wojnar , Katarzyna Kosek-Szott , Szymon Szott , Anders Jonsson , Luis Esteve , Boris Bellalta

We study the problem of scheduling packet transmissions with the aim of minimizing the energy consumption and data transmission delay of users in a wireless network in which spatial reuse of spectrum is employed. We approach this problem…

Signal Processing · Electrical Eng. & Systems 2020-06-09 Vivek S. Borkar , Shantanu Choudhary , Vaibhav Kumar Gupta , Gaurav S. Kasbekar

In restless bandits, a central agent is tasked with optimally distributing limited resources across several bandits (arms), with each arm being a Markov decision process. In this work, we generalize the traditional restless bandits problem…

Machine Learning · Computer Science 2026-02-20 Nima Akbarzadeh , Yossiri Adulyasak , Erick Delage

Reinforcement learning is an attractive approach to learn good resource allocation and scheduling policies based on data when the system model is unknown. However, the cumulative regret of most RL algorithms scales as $\tilde O(\mathsf{S}…

Machine Learning · Computer Science 2023-04-28 Nima Akbarzadeh , Aditya Mahajan