English
Related papers

Related papers: A Concentration Bound for TD(0) with Function Appr…

200 papers

The online increasing subsequence problem is a stochastic optimisation task with the objective to maximise the expected length of subsequence chosen from a random series by means of a nonanticipating decision strategy. We study the…

Probability · Mathematics 2020-01-09 Alexander Gnedin , Amirlan Seksenbayev

Given an order-$d$ tensor $\tensor A \in \R^{n \times n \times...\times n}$, we present a simple, element-wise sparsification algorithm that zeroes out all sufficiently small elements of $\tensor A$, keeps all sufficiently large elements of…

Numerical Analysis · Mathematics 2015-02-05 Nam H. Nguyen , Petros Drineas , Trac D. Tran

We study the computational complexity of approximating general constrained Markov decision processes. Our primary contribution is the design of a polynomial time $(0,\epsilon)$-additive bicriteria approximation algorithm for finding optimal…

Data Structures and Algorithms · Computer Science 2025-02-12 Jeremy McMahan

A convergence analysis is developed for the regularized Newton method for training neural networks (NNs) in the overparameterized limit. As the number of hidden units tends to infinity, the NN training dynamics converge in probability to…

Machine Learning · Computer Science 2026-05-21 Konstantin Riedl , Konstantinos Spiliopoulos , Justin Sirignano

We study stochastic approximation algorithms with Markovian noise and constant step-size $\alpha$. We develop a method based on infinitesimal generator comparisons to study the bias of the algorithm, which is the expected difference between…

Machine Learning · Statistics 2024-10-28 Sebastian Allmeier , Nicolas Gast

In this work, we study stability of distributed filtering of Markov chains with finite state space, partially observed in conditionally Gaussian noise. We consider a nonlinear filtering scheme over a Distributed Network of Agents (DNA),…

Statistics Theory · Mathematics 2016-09-27 Dionysios S. Kalogerias , Athina P. Petropulu

This work shows how exponential concentration inequalities for additive functionals of stochastic processes over a finite time interval can be derived from concentration inequalities for martingales. The approach is entirely probabilistic…

Probability · Mathematics 2020-07-14 Bob Pepin

We study the problem of learning classification functions from noiseless training samples, under the assumption that the decision boundary is of a certain regularity. We establish universal lower bounds for this estimation problem, for…

Functional Analysis · Mathematics 2021-12-28 Philipp Petersen , Felix Voigtlaender

This paper develops asymptotic theory for quantile estimation via stochastic gradient descent (SGD) with a constant learning rate. The quantile loss function is neither smooth nor strongly convex. Beyond conventional perspectives and…

Machine Learning · Statistics 2026-04-06 Ziyang Wei , Jiaqi Li , Likai Chen , Wei Biao Wu

Iterative load balancing algorithms for indivisible tokens have been studied intensively in the past. Complementing previous worst-case analyses, we study an average-case scenario where the load inputs are drawn from a fixed probability…

Distributed, Parallel, and Cluster Computing · Computer Science 2017-03-28 Leran Cai , Thomas Sauerwald

We introduce a new interpretation of the attention matrix as a discrete-time Markov chain. Our interpretation sheds light on common operations involving attention scores such as selection, summation, and averaging in a unified framework. It…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Yotam Erel , Olaf Dünkel , Rishabh Dabral , Vladislav Golyanik , Christian Theobalt , Amit H. Bermano

The work [8] established memory loss in the time-dependent (non-random) case of uniformly expanding maps of the interval. Here we find conditions under which we have convergence to the normal distribution of the appropriately scaled…

Dynamical Systems · Mathematics 2016-03-25 Peter Nandori , Domokos Szasz , Tamas Varju

We obtain moment and Gaussian bounds for general Lipschitz functions evaluated along the sample path of a Markov chain. We treat Markov chains on general (possibly unbounded) state spaces via a coupling method. If the first moment of the…

Probability · Mathematics 2010-12-08 J. -R. Chazottes , F. Redig

Using insight from numerical approximation of ODEs and the problem formulation and solution methodology of TD learning through a Galerkin relaxation, I propose a new class of TD learning algorithms. After applying the improved numerical…

Machine Learning · Computer Science 2021-04-21 Caleb Bowyer

We prove a priori bounds for solutions of stochastic reaction diffusion equations with super-linear damping in the reaction term. These bounds provide a control on the supremum of solutions on any compact space-time set which only depends…

Analysis of PDEs · Mathematics 2018-09-24 Augustin Moinat , Hendrik Weber

This paper considers the Poisson equation for general state-space Markov chains in continuous time. The main purpose of this paper is to present specific bounds for the solutions of the Poisson equation for general state-space Markov…

Probability · Mathematics 2019-09-18 Hiroyuki Masuyama

Slow mixing is the central hurdle when working with Markov chains, especially those used for Monte Carlo approximations (MCMC). In many applications, it is only of interest to estimate the stationary expectations of a small set of…

Statistics Theory · Mathematics 2016-10-04 Maxim Rabinovich , Aaditya Ramdas , Michael I. Jordan , Martin J. Wainwright

Linear TD($\lambda$) is one of the most fundamental reinforcement learning algorithms for policy evaluation. Previously, convergence rates are typically established under the assumption of linearly independent features, which does not hold…

Machine Learning · Computer Science 2025-10-15 Zixuan Xie , Xinyu Liu , Rohan Chandra , Shangtong Zhang

We study the estimation of the value function for continuous-time Markov diffusion processes using a single, discretely observed ergodic trajectory. Our work provides non-asymptotic statistical guarantees for the least-squares…

Machine Learning · Computer Science 2025-02-07 Wenlong Mou

In this paper, we are interested in investigating the perturbation bounds for the stationary distributions for discrete-time or continuous-time Markov chains on a countable state space. For discrete-time Markov chains, two new norm-wise…

Probability · Mathematics 2012-08-27 Yuanyuan Liu