中文
相关论文

相关论文: When are Kalman-filter restless bandits indexable?

200 篇论文

We study contextual bandits with finitely many actions in which the reward of each arm follows a single-index model with an arm-specific index parameter and an unknown nonparametric link function. We consider a regime in which arms…

机器学习 · 统计学 2026-03-20 Sakshi Arya , Satarupa Bhattacharjee , Bharath K. Sriperumbudur

In this article, we complement recent results on the convergence of the state estimate obtained by applying the discrete-time Kalman filter on a time-sampled continuous-time system. As the temporal discretization is refined, the estimate…

最优化与控制 · 数学 2015-12-09 Atte Aalto

A set of N independent Gaussian linear time invariant systems is observed by M sensors whose task is to provide the best possible steady-state causal minimum mean square estimate of the state of the systems, in addition to minimizing a…

最优化与控制 · 数学 2008-10-30 Jerome Le Ny , Eric Feron , Munther A. Dahleh

Whittle index policy is a powerful tool to obtain asymptotically optimal solutions for the notoriously intractable problem of restless bandits. However, finding the Whittle indices remains a difficult problem for many practical restless…

机器学习 · 计算机科学 2022-01-21 Khaled Nakhleh , Santosh Ganji , Ping-Chun Hsieh , I-Hong Hou , Srinivas Shakkottai

A more general formulation of the linear bandit problem is considered to allow for dependencies over time. Specifically, it is assumed that there exists an unknown $\mathbb{R}^d$-valued stationary $\varphi$-mixing sequence of parameters…

机器学习 · 统计学 2024-05-20 Azadeh Khaleghi

We study the Whittle index learning algorithm for restless multi-armed bandits. We consider index learning algorithm with Q-learning. We first present Q-learning algorithm with exploration policies -- epsilon-greedy, softmax,…

机器学习 · 计算机科学 2024-09-10 Vishesh Mittal , Rahul Meshram , Surya Prakash

In this paper, we consider a general observation model for restless multi-armed bandit problems. The operation of the player is based on the past observation history that is limited (partial) and error-prone due to resource constraints or…

机器学习 · 统计学 2025-12-17 Keqin Liu , Qizhen Jia

A novel reinforcement learning algorithm is introduced for multiarmed restless bandits with average reward, using the paradigms of Q-learning and Whittle index. Specifically, we leverage the structure of the Whittle index policy to reduce…

机器学习 · 计算机科学 2021-09-22 Konstantin E. Avrachenkov , Vivek S. Borkar

This article is concerned with the convergence of the state estimate obtained from the discrete time Kalman filter to the continuous time estimate as the temporal discretization is refined. We derive convergence rate estimates for different…

最优化与控制 · 数学 2015-12-10 Atte Aalto

We propose and study Collpasing Bandits, a new restless multi-armed bandit (RMAB) setting in which each arm follows a binary-state Markovian process with a special structure: when an arm is played, the state is fully observed, thus…

机器学习 · 计算机科学 2020-07-10 Aditya Mate , Jackson A. Killian , Haifeng Xu , Andrew Perrault , Milind Tambe

This work presents a notion of strong detectability for linear time varying systems affected by unknown inputs. It is shown that this notion is equivalent to detectability of an auxiliary system without unknown inputs. This allows a…

系统与控制 · 电气工程与系统科学 2021-03-24 Markus Tranninger , Richard Seeber , Juan G. Rueda-Escobedo , Martin Horn

We develop a generalization of unobserved components models that allows for a wide range of long-run dynamics by modelling the permanent component as a fractionally integrated process. The model does not require stationarity and can be cast…

计量经济学 · 经济学 2020-05-22 Tobias Hartl , Rolf Tschernig , Enzo Weber

This work studies a generalized class of restless multi-armed bandits with hidden states and allow cumulative feedback, as opposed to the conventional instantaneous feedback. We call them lazy restless bandits (LRB) as the events of…

系统与控制 · 计算机科学 2019-01-30 Kesav Kaza , Rahul Meshram , Varun Mehta , S. N. Merchant

In this paper we present a new algorithm for online (sequential) inference in Bayesian neural networks, and show its suitability for tackling contextual bandit problems. The key idea is to combine the extended Kalman filter (which locally…

机器学习 · 计算机科学 2022-05-02 Gerardo Duran-Martin , Aleyna Kara , Kevin Murphy

We present a two-armed bandit model of decision making under uncertainty where the expected return to investing in the "risky arm" increases when choosing that arm and decreases when choosing the "safe" arm. These dynamics are natural in…

最优化与控制 · 数学 2017-03-22 Roland Fryer , Philipp Harms

We study the problem of planning restless multi-armed bandits (RMABs) with multiple actions. This is a popular model for multi-agent systems with applications like multi-channel communication, monitoring and machine maintenance tasks, and…

多智能体系统 · 计算机科学 2023-03-01 Abheek Ghosh , Dheeraj Nagaraj , Manish Jain , Milind Tambe

A sensing policy for the restless multi-armed bandit problem with stationary but unknown reward distributions is proposed. The work is presented in the context of cognitive radios in which the bandit problem arises when deciding which parts…

信息论 · 计算机科学 2012-11-20 Jan Oksanen , Visa Koivunen , H. Vincent Poor

We present an analysis of ensemble Kalman inversion, based on the continuous time limit of the algorithm. The analysis of the dynamical behaviour of the ensemble allows us to establish well-posedness and convergence results for a fixed…

数值分析 · 数学 2017-08-09 Claudia Schillings , Andrew Stuart

We study the linear bandit problem that accounts for partially observable features. Without proper handling, unobserved features can lead to linear regret in the decision horizon $T$, as their influence on rewards is unknown. To tackle this…

机器学习 · 统计学 2025-08-19 Wonyoung Kim , Sungwoo Park , Garud Iyengar , Assaf Zeevi , Min-hwan Oh

Contextual bandits are a central framework for sequential decision-making, with applications ranging from recommendation systems to clinical trials. While nonparametric methods can flexibly model complex reward structures, they suffer from…

统计理论 · 数学 2026-01-01 Wanteng Ma , T. Tony Cai