English
Related papers

Related papers: A Concentration Bound for TD(0) with Function Appr…

200 papers

We study a Q learning algorithm for continuous time stochastic control problems. The proposed algorithm uses the sampled state process by discretizing the state and control action spaces under piece-wise constant control processes. We show…

Optimization and Control · Mathematics 2023-03-10 Erhan Bayraktar , Ali Devran Kara

Temporal-difference learning with gradient correction (TDC) is a two time-scale algorithm for policy evaluation in reinforcement learning. This algorithm was initially proposed with linear function approximation, and was later extended to…

Machine Learning · Computer Science 2021-10-29 Yue Wang , Shaofeng Zou , Yi Zhou

We consider a Markov chain $(M_{n})_{n\ge 0}$ on the set $\mathbb{N}_{0}$ of nonnegative integers which is eventually decreasing, i.e. $\mathbb{P}\{M_{n+1}<M_{n}|M_{n}\ge a\}=1$ for some $a\in\mathbb{N}$ and all $n\ge 0$. We are interested…

Probability · Mathematics 2015-09-08 Gerold Alsmeyer , Alexander Marynych

Machine learning models with inputs in a Euclidean space $\mathbb{R}^d$, when implemented on digital computers, generalize, and their generalization gap converges to $0$ at a rate of $c/N^{1/2}$ concerning the sample size $N$. However, the…

Machine Learning · Computer Science 2026-05-14 Anastasis Kratsios , A. Martina Neuman , Gudmund Pammer

Motivated by queues with many servers, we study Brownian steady-state approximations for continuous time Markov chains (CTMCs). Our approximations are based on diffusion models (rather than a diffusion limit) whose steady-state, we prove,…

Probability · Mathematics 2014-09-12 Itai Gurvich

The solution to Poisson's equation arise in many Markov chain and Markov jump process settings, including that of the central limit theorem, value functions for average reward Markov decision processes, and within the gradient formula for…

Probability · Mathematics 2024-01-30 Saied Mahdian , Peter W. Glynn , Yuanyuan Liu

We study the problem of learning general (i.e., not necessarily homogeneous) halfspaces with Random Classification Noise under the Gaussian distribution. We establish nearly-matching algorithmic and Statistical Query (SQ) lower bound…

Machine Learning · Computer Science 2023-07-18 Ilias Diakonikolas , Jelena Diakonikolas , Daniel M. Kane , Puqian Wang , Nikos Zarifis

We consider the basic problem of learning Single-Index Models with respect to the square loss under the Gaussian distribution in the presence of adversarial label noise. Our main contribution is the first computationally efficient algorithm…

Machine Learning · Computer Science 2025-08-07 Puqian Wang , Nikos Zarifis , Ilias Diakonikolas , Jelena Diakonikolas

Poisson's equation is fundamental to the study of Markov chains, and arises in connection with martingale representations and central limit theorems for additive functionals, perturbation theory for stationary distributions, and average…

Probability · Mathematics 2025-04-03 Peter W. Glynn , Na Lin , Yuanyuan Liu

This paper is dedicated to the investigation of a new numerical method to approximate the optimal stopping problem for a discrete-time continuous state space Markov chain under partial observations. It is based on a two-step discretization…

Optimization and Control · Mathematics 2016-02-16 Benoîte de Saporta , François Dufour , Christophe Nivot

We provide a brief tutorial on the use of concentration inequalities as they apply to system identification of state-space parameters of linear time invariant systems, with a focus on the fully observed setting. We draw upon tools from the…

Optimization and Control · Mathematics 2019-08-30 Nikolai Matni , Stephen Tu

Estimates are constructed for the deviation of the concentration functions of sums of independent random variables with finite variances from the folded normal distribution function without any assumptions concerning the existence of the…

Probability · Mathematics 2016-08-11 V. Yu. Korolev , A. V. Dorofeeva

We consider the problem of balancing load items (tokens) in networks. Starting with an arbitrary load distribution, we allow nodes to exchange tokens with their neighbors in each round. The goal is to achieve a distribution where all nodes…

Discrete Mathematics · Computer Science 2015-03-20 Thomas Sauerwald , He Sun

We obtain non asymptotic concentration bounds for two kinds of stochastic approximations. We first consider the deviations between the expectation of a given function of the Euler scheme of some diffusion process at a fixed deterministic…

Probability · Mathematics 2012-12-12 Noufel Frikha , Stephane Menozzi

We consider a standard distributed optimisation setting where $N$ machines, each holding a $d$-dimensional function $f_i$, aim to jointly minimise the sum of the functions $\sum_{i = 1}^N f_i (x)$. This problem arises naturally in…

Machine Learning · Computer Science 2021-12-08 Dan Alistarh , Janne H. Korhonen

This paper is concerned with the hard thresholding operator which sets all but the $k$ largest absolute elements of a vector to zero. We establish a {\em tight} bound to quantitatively characterize the deviation of the thresholded solution…

Machine Learning · Statistics 2020-08-12 Jie Shen , Ping Li

Diffusion models have achieved huge empirical success in data generation tasks. Recently, some efforts have been made to adapt the framework of diffusion models to discrete state space, providing a more natural approach for modeling…

Machine Learning · Statistics 2024-02-15 Hongrui Chen , Lexing Ying

We prove an apparently novel concentration of measure result for Markov tree processes. The bound we derive reduces to the known bounds for Markov processes when the tree is a chain, thus strictly generalizing the known Markov process…

Probability · Mathematics 2007-05-23 Leonid Kontorovich

We study the complexity of learning and approximation of self-bounding functions over the uniform distribution on the Boolean hypercube ${0,1}^n$. Informally, a function $f:{0,1}^n \rightarrow \mathbb{R}$ is self-bounding if for every $x…

Machine Learning · Computer Science 2019-06-04 Vitaly Feldman , Pravesh Kothari , Jan Vondrák

We study two time-scale linear stochastic approximation algorithms, which can be used to model well-known reinforcement learning algorithms such as GTD, GTD2, and TDC. We present finite-time performance bounds for the case where the…

Machine Learning · Computer Science 2019-07-16 Harsh Gupta , R. Srikant , Lei Ying
‹ Prev 1 4 5 6 7 8 10 Next ›