English
Related papers

Related papers: When are Kalman-filter restless bandits indexable?

200 papers

We study contextual bandits with finitely many actions in which the reward of each arm follows a single-index model with an arm-specific index parameter and an unknown nonparametric link function. We consider a regime in which arms…

Machine Learning · Statistics 2026-03-20 Sakshi Arya , Satarupa Bhattacharjee , Bharath K. Sriperumbudur

In this article, we complement recent results on the convergence of the state estimate obtained by applying the discrete-time Kalman filter on a time-sampled continuous-time system. As the temporal discretization is refined, the estimate…

Optimization and Control · Mathematics 2015-12-09 Atte Aalto

A set of N independent Gaussian linear time invariant systems is observed by M sensors whose task is to provide the best possible steady-state causal minimum mean square estimate of the state of the systems, in addition to minimizing a…

Optimization and Control · Mathematics 2008-10-30 Jerome Le Ny , Eric Feron , Munther A. Dahleh

Whittle index policy is a powerful tool to obtain asymptotically optimal solutions for the notoriously intractable problem of restless bandits. However, finding the Whittle indices remains a difficult problem for many practical restless…

Machine Learning · Computer Science 2022-01-21 Khaled Nakhleh , Santosh Ganji , Ping-Chun Hsieh , I-Hong Hou , Srinivas Shakkottai

A more general formulation of the linear bandit problem is considered to allow for dependencies over time. Specifically, it is assumed that there exists an unknown $\mathbb{R}^d$-valued stationary $\varphi$-mixing sequence of parameters…

Machine Learning · Statistics 2024-05-20 Azadeh Khaleghi

We study the Whittle index learning algorithm for restless multi-armed bandits. We consider index learning algorithm with Q-learning. We first present Q-learning algorithm with exploration policies -- epsilon-greedy, softmax,…

Machine Learning · Computer Science 2024-09-10 Vishesh Mittal , Rahul Meshram , Surya Prakash

In this paper, we consider a general observation model for restless multi-armed bandit problems. The operation of the player is based on the past observation history that is limited (partial) and error-prone due to resource constraints or…

Machine Learning · Statistics 2025-12-17 Keqin Liu , Qizhen Jia

A novel reinforcement learning algorithm is introduced for multiarmed restless bandits with average reward, using the paradigms of Q-learning and Whittle index. Specifically, we leverage the structure of the Whittle index policy to reduce…

Machine Learning · Computer Science 2021-09-22 Konstantin E. Avrachenkov , Vivek S. Borkar

This article is concerned with the convergence of the state estimate obtained from the discrete time Kalman filter to the continuous time estimate as the temporal discretization is refined. We derive convergence rate estimates for different…

Optimization and Control · Mathematics 2015-12-10 Atte Aalto

We propose and study Collpasing Bandits, a new restless multi-armed bandit (RMAB) setting in which each arm follows a binary-state Markovian process with a special structure: when an arm is played, the state is fully observed, thus…

Machine Learning · Computer Science 2020-07-10 Aditya Mate , Jackson A. Killian , Haifeng Xu , Andrew Perrault , Milind Tambe

This work presents a notion of strong detectability for linear time varying systems affected by unknown inputs. It is shown that this notion is equivalent to detectability of an auxiliary system without unknown inputs. This allows a…

Systems and Control · Electrical Eng. & Systems 2021-03-24 Markus Tranninger , Richard Seeber , Juan G. Rueda-Escobedo , Martin Horn

We develop a generalization of unobserved components models that allows for a wide range of long-run dynamics by modelling the permanent component as a fractionally integrated process. The model does not require stationarity and can be cast…

Econometrics · Economics 2020-05-22 Tobias Hartl , Rolf Tschernig , Enzo Weber

This work studies a generalized class of restless multi-armed bandits with hidden states and allow cumulative feedback, as opposed to the conventional instantaneous feedback. We call them lazy restless bandits (LRB) as the events of…

Systems and Control · Computer Science 2019-01-30 Kesav Kaza , Rahul Meshram , Varun Mehta , S. N. Merchant

In this paper we present a new algorithm for online (sequential) inference in Bayesian neural networks, and show its suitability for tackling contextual bandit problems. The key idea is to combine the extended Kalman filter (which locally…

Machine Learning · Computer Science 2022-05-02 Gerardo Duran-Martin , Aleyna Kara , Kevin Murphy

We present a two-armed bandit model of decision making under uncertainty where the expected return to investing in the "risky arm" increases when choosing that arm and decreases when choosing the "safe" arm. These dynamics are natural in…

Optimization and Control · Mathematics 2017-03-22 Roland Fryer , Philipp Harms

We study the problem of planning restless multi-armed bandits (RMABs) with multiple actions. This is a popular model for multi-agent systems with applications like multi-channel communication, monitoring and machine maintenance tasks, and…

Multiagent Systems · Computer Science 2023-03-01 Abheek Ghosh , Dheeraj Nagaraj , Manish Jain , Milind Tambe

A sensing policy for the restless multi-armed bandit problem with stationary but unknown reward distributions is proposed. The work is presented in the context of cognitive radios in which the bandit problem arises when deciding which parts…

Information Theory · Computer Science 2012-11-20 Jan Oksanen , Visa Koivunen , H. Vincent Poor

We present an analysis of ensemble Kalman inversion, based on the continuous time limit of the algorithm. The analysis of the dynamical behaviour of the ensemble allows us to establish well-posedness and convergence results for a fixed…

Numerical Analysis · Mathematics 2017-08-09 Claudia Schillings , Andrew Stuart

We study the linear bandit problem that accounts for partially observable features. Without proper handling, unobserved features can lead to linear regret in the decision horizon $T$, as their influence on rewards is unknown. To tackle this…

Machine Learning · Statistics 2025-08-19 Wonyoung Kim , Sungwoo Park , Garud Iyengar , Assaf Zeevi , Min-hwan Oh

Contextual bandits are a central framework for sequential decision-making, with applications ranging from recommendation systems to clinical trials. While nonparametric methods can flexibly model complex reward structures, they suffer from…

Statistics Theory · Mathematics 2026-01-01 Wanteng Ma , T. Tony Cai