Related papers: On the Stability of Random Matrix Product with Mar…
We are interested in understanding stability (almost sure boundedness) of stochastic approximation algorithms (SAs) driven by a `controlled Markov' process. Analyzing this class of algorithms is important, since many reinforcement learning…
We consider the dynamics of a linear stochastic approximation algorithm driven by Markovian noise, and derive finite-time bounds on the moments of the error, i.e., deviation of the output of the algorithm from the equilibrium point of an…
In this paper we consider the problem of obtaining sharp bounds for the performance of temporal difference (TD) methods with linear function approximation for policy evaluation in discounted Markov decision processes. We show that a simple…
It is known that state-dependent, multi-step Lyapunov bounds lead to greatly simplified verification theorems for stability for large classes of Markov chain models. This is one component of the "fluid model" approach to stability of…
The paper deals with the convergence properties of the products of random (row-)stochastic matrices. The limiting behavior of such products is studied from a dynamical system point of view. In particular, by appropriately defining a dynamic…
Motivated by engineering applications such as resource allocation in networks and inventory systems, we consider average-reward Reinforcement Learning with unbounded state space and reward function. Recent works studied this problem in the…
We develop a practical approach to establish the stability, that is, the recurrence in a given set, of a large class of controlled Markov chains. These processes arise in various areas of applied science and encompass important numerical…
We study stochastic approximation procedures for approximately solving a $d$-dimensional linear fixed point equation based on observing a trajectory of length $n$ from an ergodic Markov chain. We first exhibit a non-asymptotic bound of the…
This paper studies the finite-time stability and stabilization of linear discrete time-varying stochastic systems with multiplicative noise. Firstly, necessary and sufficient conditions for finite-time stability are presented via state…
Temporal-difference (TD) learning is widely regarded as one of the most popular algorithms in reinforcement learning (RL). Despite its widespread use, it has only been recently that researchers have begun to actively study its finite time…
We present for the first time an asymptotic convergence analysis of two time-scale stochastic approximation driven by "controlled" Markov noise. In particular, the faster and slower recursions have non-additive controlled Markov noise…
We establish novel and general high-dimensional concentration inequalities and Berry-Esseen bounds for vector-valued martingales induced by Markov chains. We apply these results to analyze the performance of the Temporal Difference (TD)…
The problem of p-th moment stability for time-varying stochastic time-delay systems with Markovian switching is investigated in this paper. Some novel stability criteria are obtained by applying the generalized Razumikhin and Krasovskii…
This work is concerned with the stability properties of linear stochastic differential equations with random (drift and diffusion) coefficient matrices, and the stability of a corresponding random transition matrix (or exponential…
Temporal difference learning (TD) is a simple iterative algorithm used to estimate the value function corresponding to a given policy in a Markov decision process. Although TD is one of the most widely used algorithms in reinforcement…
This paper aims to develop the stability theory for singular stochastic Markov jump systems with state-dependent noise, including both continuous- and discrete-time cases. The sufficient conditions for the existence and uniqueness of a…
One of the most basic problems in reinforcement learning (RL) is policy evaluation: estimating the long-term return, i.e., value function, corresponding to a given fixed policy. The celebrated Temporal Difference (TD) learning algorithm…
This paper continues the discussion on the stability of time-inhomogeneous Markov chains. In particular, this paper defines a time-inhomogeneous, discrete-time Markov chain governed by a continuous evolution in the appropriate martrix…
We study the singular values and Lyapunov exponents of non-stationary random matrix products subject to small, absolutely continuous, additive noise. Consider a fixed sequence of matrices of bounded norm. Independently perturb the matrices…
We study the policy evaluation problem in multi-agent reinforcement learning, modeled by a Markov decision process. In this problem, the agents operate in a common environment under a fixed control policy, working together to discover the…