English
Related papers

Related papers: A fast-pivoting algorithm for Whittle's restless b…

200 papers

Scheduling in multi-channel wireless communication system presents formidable challenges in effectively allocating resources. To address these challenges, we investigate a multi-resource restless matching bandit (MR-RMB) model for…

Machine Learning · Computer Science 2024-08-21 Nida Zamir , I-Hong Hou

We consider the problem of maximizing the expected average reward obtained over an infinite time horizon by $n$ weakly coupled Markov decision processes. Our setup is a substantial generalization of the multi-armed restless bandit problem…

Optimization and Control · Mathematics 2026-04-01 Diego Goldsztajn , Konstantin Avrachenkov

We study the problem of identifying the best arm in a multi-armed bandit environment when each arm is a time-homogeneous and ergodic discrete-time Markov process on a common, finite state space. The state evolution on each arm is governed…

Machine Learning · Statistics 2022-03-30 P. N. Karthik , Kota Srinivas Reddy , Vincent Y. F. Tan

Non-stationary parametric bandits have attracted much attention recently. There are three principled ways to deal with non-stationarity, including sliding-window, weighted, and restart strategies. As many non-stationary environments exhibit…

Machine Learning · Computer Science 2026-01-06 Jing Wang , Peng Zhao , Zhi-Hua Zhou

We address the problem of opportunistic multiuser scheduling in downlink networks with Markov-modeled outage channels. We consider the scenario in which the scheduler does not have full knowledge of the channel state information, but…

Networking and Internet Architecture · Computer Science 2011-12-08 Wenzhuo Ouyang , Sugumar Murugesan , Atilla Eryilmaz , Ness B. Shroff

In this paper, we consider a general observation model for restless multi-armed bandit problems. The operation of the player is based on the past observation history that is limited (partial) and error-prone due to resource constraints or…

Machine Learning · Statistics 2025-12-17 Keqin Liu , Qizhen Jia

Contextual bandits are a central framework for sequential decision-making, with applications ranging from recommendation systems to clinical trials. While nonparametric methods can flexibly model complex reward structures, they suffer from…

Statistics Theory · Mathematics 2026-01-01 Wanteng Ma , T. Tony Cai

In this paper we address the problem of allocating the efforts of a collection of repairmen to a number of deteriorating machines in order to reduce operation costs and to mitigate the cost (and likelihood) of unexpected failures.…

Discrete Mathematics · Computer Science 2024-01-26 Diego Ruiz-Hernandez , Jesús María Pinar-Pérez , David Delgado-Gómez

Non-stationary parametric bandits have attracted much attention recently. There are three principled ways to deal with non-stationarity, including sliding-window, weighted, and restart strategies. As many non-stationary environments exhibit…

Machine Learning · Computer Science 2023-06-08 Jing Wang , Peng Zhao , Zhi-Hua Zhou

We study a system with finitely many groups of multi-action bandit processes, each of which is a Markov decision process (MDP) with finite state and action spaces and potentially different transition matrices when taking different actions.…

Optimization and Control · Mathematics 2024-12-05 Jing Fu , Bill Moran , José Niño-Mora

The Gittins index is a tool that optimally solves a variety of decision-making problems involving uncertainty, including multi-armed bandit problems, minimizing mean latency in queues, and search problems like the Pandora's box model.…

Optimization and Control · Mathematics 2025-08-05 Ziv Scully , Alexander Terenin

The Greedy algorithm is the simplest heuristic in sequential decision problem that carelessly takes the locally optimal choice at each round, disregarding any advantages of exploring and/or information gathering. Theoretically, it is known…

Machine Learning · Computer Science 2021-01-05 Matthieu Jedor , Jonathan Louëdec , Vianney Perchet

Restless multi-armed bandits (RMABs) have been widely utilized to address resource allocation problems with Markov reward processes (MRPs). Existing works often assume that the dynamics of MRPs are known prior, which makes the RMAB problem…

Machine Learning · Computer Science 2024-06-13 Jingwen Tong , Xinran Li , Liqun Fu , Jun Zhang , Khaled B. Letaief

This paper is in the field of stochastic Multi-Armed Bandits (MABs), i.e., those sequential selection techniques able to learn online using only the feedback given by the chosen option (a.k.a. arm). We study a particular case of the rested…

Machine Learning · Computer Science 2022-12-08 Alberto Maria Metelli , Francesco Trovò , Matteo Pirola , Marcello Restelli

In this paper, we explore the use of multi-armed bandit online learning techniques to solve distributed resource selection problems. As an example, we focus on the problem of network selection. Mobile devices often have several wireless…

Computer Science and Game Theory · Computer Science 2018-05-15 Anuja Meetoo Appavoo , Seth Gilbert , Kian-Lee Tan

Restless multi-armed bandits (RMAB) play a central role in modeling sequential decision making problems under an instantaneous activation constraint that at most B arms can be activated at any decision epoch. Each restless arm is endowed…

Machine Learning · Computer Science 2024-05-03 Guojun Xiong , Jian Li

Restless multi-armed bandits (RMABs) have been highly successful in optimizing sequential resource allocation across many domains. However, in many practical settings with highly scarce resources, where each agent can only receive at most…

Multiagent Systems · Computer Science 2025-01-13 Guojun Xiong , Haichuan Wang , Yuqi Pan , Saptarshi Mandal , Sanket Shah , Niclas Boehmer , Milind Tambe

We consider a policy gradient algorithm applied to a finite-arm bandit problem with Bernoulli rewards. We allow learning rates to depend on the current state of the algorithm, rather than use a deterministic time-decreasing learning rate.…

Machine Learning · Computer Science 2021-09-24 Denis Denisov , Neil Walton

We study the $\textit{single-index bandit}$ problem, where rewards depend on an unknown one-dimensional projection of high-dimensional contexts through an unknown reward function. This model extends linear and generalized linear bandits to…

Machine Learning · Statistics 2026-05-12 Devdan Dey , Sujoy Bhore , Avishek Ghosh

Restless multi-armed bandits (RMAB) extend multi-armed bandits so pulling an arm impacts future states. Despite the success of RMABs, a key limiting assumption is the separability of rewards into a sum across arms. We address this…

Machine Learning · Computer Science 2024-06-11 Naveen Raman , Zheyuan Ryan Shi , Fei Fang
‹ Prev 1 4 5 6 7 8 10 Next ›