中文
相关论文

相关论文: On the Stability of Random Matrix Product with Mar…

200 篇论文

We are interested in understanding stability (almost sure boundedness) of stochastic approximation algorithms (SAs) driven by a `controlled Markov' process. Analyzing this class of algorithms is important, since many reinforcement learning…

系统与控制 · 计算机科学 2018-05-18 Arunselvan Ramaswamy , Shalabh Bhatnagar

We consider the dynamics of a linear stochastic approximation algorithm driven by Markovian noise, and derive finite-time bounds on the moments of the error, i.e., deviation of the output of the algorithm from the equilibrium point of an…

机器学习 · 计算机科学 2019-03-11 R. Srikant , Lei Ying

In this paper we consider the problem of obtaining sharp bounds for the performance of temporal difference (TD) methods with linear function approximation for policy evaluation in discounted Markov decision processes. We show that a simple…

机器学习 · 统计学 2024-06-18 Sergey Samsonov , Daniil Tiapkin , Alexey Naumov , Eric Moulines

It is known that state-dependent, multi-step Lyapunov bounds lead to greatly simplified verification theorems for stability for large classes of Markov chain models. This is one component of the "fluid model" approach to stability of…

最优化与控制 · 数学 2012-05-18 Serdar Yüksel , Sean P. Meyn

The paper deals with the convergence properties of the products of random (row-)stochastic matrices. The limiting behavior of such products is studied from a dynamical system point of view. In particular, by appropriately defining a dynamic…

概率论 · 数学 2013-01-15 Behrouz Touri , Angelia Nedich

Motivated by engineering applications such as resource allocation in networks and inventory systems, we consider average-reward Reinforcement Learning with unbounded state space and reward function. Recent works studied this problem in the…

机器学习 · 计算机科学 2025-11-10 Shaan Ul Haque , Siva Theja Maguluri

We develop a practical approach to establish the stability, that is, the recurrence in a given set, of a large class of controlled Markov chains. These processes arise in various areas of applied science and encompass important numerical…

统计理论 · 数学 2015-02-02 Christophe Andrieu , Vladislav B. Tadić , Matti Vihola

We study stochastic approximation procedures for approximately solving a $d$-dimensional linear fixed point equation based on observing a trajectory of length $n$ from an ergodic Markov chain. We first exhibit a non-asymptotic bound of the…

最优化与控制 · 数学 2024-05-14 Wenlong Mou , Ashwin Pananjady , Martin J. Wainwright , Peter L. Bartlett

This paper studies the finite-time stability and stabilization of linear discrete time-varying stochastic systems with multiplicative noise. Firstly, necessary and sufficient conditions for finite-time stability are presented via state…

最优化与控制 · 数学 2018-06-25 Tianliang Zhang , Feiqi Deng , Weihai Zhang

Temporal-difference (TD) learning is widely regarded as one of the most popular algorithms in reinforcement learning (RL). Despite its widespread use, it has only been recently that researchers have begun to actively study its finite time…

机器学习 · 计算机科学 2025-04-16 Han-Dong Lim , Donghwan Lee

We present for the first time an asymptotic convergence analysis of two time-scale stochastic approximation driven by "controlled" Markov noise. In particular, the faster and slower recursions have non-additive controlled Markov noise…

机器学习 · 计算机科学 2020-12-03 Prasenjit Karmakar

We establish novel and general high-dimensional concentration inequalities and Berry-Esseen bounds for vector-valued martingales induced by Markov chains. We apply these results to analyze the performance of the Temporal Difference (TD)…

机器学习 · 统计学 2026-05-22 Weichen Wu , Yuting Wei , Alessandro Rinaldo

The problem of p-th moment stability for time-varying stochastic time-delay systems with Markovian switching is investigated in this paper. Some novel stability criteria are obtained by applying the generalized Razumikhin and Krasovskii…

动力系统 · 数学 2016-07-11 Bin Zhou , Weiwei Luo

This work is concerned with the stability properties of linear stochastic differential equations with random (drift and diffusion) coefficient matrices, and the stability of a corresponding random transition matrix (or exponential…

概率论 · 数学 2019-05-02 Adrian N. Bishop , Pierre Del Moral

Temporal difference learning (TD) is a simple iterative algorithm used to estimate the value function corresponding to a given policy in a Markov decision process. Although TD is one of the most widely used algorithms in reinforcement…

机器学习 · 计算机科学 2018-11-07 Jalaj Bhandari , Daniel Russo , Raghav Singal

This paper aims to develop the stability theory for singular stochastic Markov jump systems with state-dependent noise, including both continuous- and discrete-time cases. The sufficient conditions for the existence and uniqueness of a…

最优化与控制 · 数学 2015-09-04 Yong Zhao , Weihai Zhang

One of the most basic problems in reinforcement learning (RL) is policy evaluation: estimating the long-term return, i.e., value function, corresponding to a given fixed policy. The celebrated Temporal Difference (TD) learning algorithm…

机器学习 · 计算机科学 2025-02-10 Sreejeet Maity , Aritra Mitra

This paper continues the discussion on the stability of time-inhomogeneous Markov chains. In particular, this paper defines a time-inhomogeneous, discrete-time Markov chain governed by a continuous evolution in the appropriate martrix…

概率论 · 数学 2015-07-23 Kyle Bradford

We study the singular values and Lyapunov exponents of non-stationary random matrix products subject to small, absolutely continuous, additive noise. Consider a fixed sequence of matrices of bounded norm. Independently perturb the matrices…

概率论 · 数学 2025-12-22 Sam Bednarski , Jonathan DeWitt , Anthony Quas

We study the policy evaluation problem in multi-agent reinforcement learning, modeled by a Markov decision process. In this problem, the agents operate in a common environment under a fixed control policy, working together to discover the…

最优化与控制 · 数学 2020-01-13 Thinh T. Doan , Siva Theja Maguluri , Justin Romberg
‹ 上一页 1 2 3 10 下一页 ›