English
Related papers

Related papers: Gaussian Approximation for Asynchronous Q-learning

200 papers

Consider a Markov decision process (MDP) that admits a set of state-action features, which can linearly express the process's probabilistic transition model. We propose a parametric Q-learning algorithm that finds an approximate-optimal…

Machine Learning · Computer Science 2019-06-07 Lin F. Yang , Mengdi Wang

We revisit the unified two-timescale Q-learning algorithm as initially introduced by Angiuli et al. \cite{angiuli2022unified}. This algorithm demonstrates efficacy in solving mean field game (MFG) and mean field control (MFC) problems,…

Optimization and Control · Mathematics 2024-05-30 Jing An , Jianfeng Lu , Yue Wu , Yang Xiang

Machine learning practitioners invest significant manual and computational resources in finding suitable learning rates for optimization algorithms. We provide a probabilistic motivation, in terms of Gaussian inference, for popular…

Machine Learning · Computer Science 2021-02-23 Filip de Roos , Carl Jidling , Adrian Wills , Thomas Schön , Philipp Hennig

Large-scale multi-agent systems are often deployed across wide geographic areas, where agents interact with heterogeneous environments. There is an emerging interest in understanding the role of heterogeneity in the performance of the…

Machine Learning · Computer Science 2026-05-18 Leo Muxing Wang , Pengkun Yang , Lili Su

Approximating the solution of the nonlinear filtering problem with Gaussian mixtures has been a very popular method since the 1970s. However, the vast majority of such approximations are introduced in an ad-hoc manner without theoretical…

Probability · Mathematics 2014-01-28 Dan Crisan , Kai Li

We study the problem of learning general (i.e., not necessarily homogeneous) halfspaces with Random Classification Noise under the Gaussian distribution. We establish nearly-matching algorithmic and Statistical Query (SQ) lower bound…

Machine Learning · Computer Science 2023-07-18 Ilias Diakonikolas , Jelena Diakonikolas , Daniel M. Kane , Puqian Wang , Nikos Zarifis

Q-learning is a stochastic approximation version of the classic value iteration. The literature has established that Q-learning suffers from both maximization bias and slower convergence. Recently, multi-step algorithms have shown practical…

Machine Learning · Computer Science 2024-07-03 Antony Vijesh , Shreyas S R

We study the rate of convergence of linear two-time-scale stochastic approximation methods. We consider two-time-scale linear iterations driven by i.i.d. noise, prove some results on their asymptotic covariance and establish asymptotic…

Probability · Mathematics 2009-09-29 Vijay R. Konda , John N. Tsitsiklis

In this paper, we give a quadratic Goldreich-Levin algorithm that is close to optimal in the following ways. Given a bounded function $f$ on the Boolean hypercube $\mathbb{F}_2^n$ and any $\varepsilon>0$, the algorithm returns a quadratic…

Computational Complexity · Computer Science 2025-05-20 Jop Briët , Davi Castro-Silva

In this paper, as a study of reinforcement learning, we converge the Q function to unbounded rewards such as Gaussian distribution. From the central limit theorem, in some real-world applications it is natural to assume that rewards follow…

Optimization and Control · Mathematics 2021-09-14 Konatsu Miyamoto , Masaya Suzuki , Yuma Kigami , Kodai Satake

In this paper, we obtain the Berry-Esseen bound for multivariate normal approximation for the Polyak-Ruppert averaged iterates of the linear stochastic approximation (LSA) algorithm with decreasing step size. Moreover, we prove the…

Machine Learning · Statistics 2025-02-04 Sergey Samsonov , Eric Moulines , Qi-Man Shao , Zhuo-Song Zhang , Alexey Naumov

We study the problem of learning hierarchical polynomials over the standard Gaussian distribution with three-layer neural networks. We specifically consider target functions of the form $h = g \circ p$ where $p : \mathbb{R}^d \rightarrow…

Machine Learning · Computer Science 2023-11-27 Zihao Wang , Eshaan Nichani , Jason D. Lee

This paper derives central limit and bootstrap theorems for probabilities that sums of centered high-dimensional random vectors hit hyperrectangles and sparsely convex sets. Specifically, we derive Gaussian and bootstrap approximations for…

Statistics Theory · Mathematics 2016-03-09 Victor Chernozhukov , Denis Chetverikov , Kengo Kato

In this article we establish new central limit theorems for Ruppert-Polyak averaged stochastic gradient descent schemes. Compared to previous work we do not assume that convergence occurs to an isolated attractor but instead allow…

Probability · Mathematics 2019-12-20 Steffen Dereich , Sebastian Kassing

We analyze the Bayesian regret of the Gaussian process posterior sampling reinforcement learning (GP-PSRL) algorithm. Posterior sampling is an effective heuristic for decision-making under uncertainty that has been used to develop…

Machine Learning · Statistics 2026-03-10 Hamish Flynn , Joe Watson , Ingmar Posner , Jan Peters

Although Q-learning is one of the most successful algorithms for finding the best action-value function (and thus the optimal policy) in reinforcement learning, its implementation often suffers from large overestimation of Q-function values…

Machine Learning · Computer Science 2020-10-13 Huaqing Xiong , Lin Zhao , Yingbin Liang , Wei Zhang

This paper is concerned with the design of algorithms based on systems of interacting particles to represent, approximate, and learn the optimal control law for reinforcement learning (RL). The primary contribution is that convergence rates…

Systems and Control · Electrical Eng. & Systems 2025-10-21 Anant A Joshi , Heng-Sheng Chang , Amirhossein Taghvaei , Prashant G Mehta , Sean P. Meyn

We present a new approach to the bootstrap for chains of infinite order taking values on a finite alphabet. It is based on a sequential Bootstrap Central Limit Theorem for the sequence of canonical Markov approximations of the chain of…

Probability · Mathematics 2007-05-23 P. Collet , D. Duarte , A. Galves

The quantum central limit theorem for bosonic quantum systems states that the sequence of states $\rho^{\boxplus n}$ obtained from the $n$-fold convolution of a centered quantum state $\rho$ converges to a quantum Gaussian state $\rho_G$…

Quantum Physics · Physics 2025-08-01 Salman Beigi , Hami Mehrabi

Despite the sustained popularity of Q-learning as a practical tool for policy determination, a majority of relevant theoretical literature deals with either constant ($\eta_{t}\equiv \eta$) or polynomially decaying ($\eta_{t} = \eta…

Machine Learning · Statistics 2026-04-07 Soham Bonnerjee , Zhipeng Lou , Wei Biao Wu