English
Related papers

Related papers: Heavy-Tailed and Long-Range Dependent Noise in Sto…

200 papers

Many machine learning and optimization algorithms can be cast as instances of stochastic approximation (SA). The convergence rate of these algorithms is known to be slow, with the optimal mean squared error (MSE) of order $O(n^{-1})$. In…

Optimization and Control · Mathematics 2024-09-13 Caio Kalil Lauand , Sean Meyn

Heavy-tail phenomena in stochastic gradient descent (SGD) have been reported in several empirical studies. Experimental evidence in previous works suggests a strong interplay between the heaviness of the tails and generalization behavior of…

Machine Learning · Statistics 2023-01-31 Anant Raj , Lingjiong Zhu , Mert Gürbüzbalaban , Umut Şimşekli

In this work, we study the asymptotic randomness of an algorithmic estimator of the saddle point of a globally convex-concave and locally strongly-convex strongly-concave objective. Specifically, we show that the averaged iterates of a…

Optimization and Control · Mathematics 2023-11-07 Abhishek Roy , Yi-An Ma

Temporal difference (TD) learning is a cornerstone reinforcement learning (RL) method for policy evaluation, where the goal is to estimate the value function of a Markov decision process under a fixed policy. While a substantial body of…

Machine Learning · Computer Science 2026-02-02 Donghwan Lee , Do Wan Kim

We study a stochastic optimization problem in which the sampling distribution depends on the decision variable, and the available samples are generated through an iterate-dependent Markov chain. Such settings arise naturally in problems…

Optimization and Control · Mathematics 2026-05-18 Anik Kumar Paul , Shalabh Bhatnagar

In this paper, we propose practical normalized stochastic first-order methods with Polyak momentum, multi-extrapolated momentum, and recursive momentum for solving unconstrained optimization problems. These methods employ dynamically…

Optimization and Control · Mathematics 2026-02-12 Chuan He , Zhaosong Lu , Defeng Sun , Zhanwang Deng

We analyze neural scaling laws in a solvable model of last-layer fine-tuning where targets have intrinsic, instance-heterogeneous difficulty. In our Latent Instance Difficulty (LID) model, each input's target variance is governed by a…

Machine Learning · Computer Science 2026-01-08 Noam Levi

We develop a complete and rigorous mathematical framework for the analysis of stochastic neural field equations under the influence of spatially extended additive noise. By comparing a solution to a fixed deterministic front profile it is…

Probability · Mathematics 2019-02-11 Jennifer Krüger , Wilhelm Stannat

Stochastic Gradient Descent (SGD) is the workhorse algorithm of deep learning technology. At each step of the training phase, a mini batch of samples is drawn from the training dataset and the weights of the neural network are adjusted…

Disordered Systems and Neural Networks · Physics 2022-09-07 Francesca Mignacco , Pierfrancesco Urbani

It has repeatedly been observed that loss minimization by stochastic gradient descent (SGD) leads to heavy-tailed distributions of neural network parameters. Here, we analyze a continuous diffusion approximation of SGD, called homogenized…

Machine Learning · Statistics 2024-02-05 Zhe Jiao , Martin Keller-Ressel

The paper introduces structured machine learning regressions for heavy-tailed dependent panel data potentially sampled at different frequencies. We focus on the sparse-group LASSO regularization. This type of regularization can take…

Econometrics · Economics 2021-11-23 Andrii Babii , Ryan T. Ball , Eric Ghysels , Jonas Striaukas

This paper presents uniform-in-time finite-sample bounds for regularized linear regression with vector-valued outputs and conditionally zero-mean subgaussian noise. By revisiting classical self-normalized martingale arguments, we obtain…

Statistics Theory · Mathematics 2026-03-20 Léo Simpson , Katrin Baumgärtner , Johannes Köhler , Moritz Diehl

Reinforcement learning is a framework for interactive decision-making with incentives sequentially revealed across time without a system dynamics model. Due to its scaling to continuous spaces, we focus on policy search where one…

Machine Learning · Computer Science 2023-01-04 Amrit Singh Bedi , Anjaly Parayil , Junyu Zhang , Mengdi Wang , Alec Koppel

We investigate the high-dimensional properties of robust regression estimators in the presence of heavy-tailed contamination of both the covariates and response functions. In particular, we provide a sharp asymptotic characterisation of…

Statistics Theory · Mathematics 2024-06-03 Urte Adomaityte , Leonardo Defilippis , Bruno Loureiro , Gabriele Sicuro

We study a stochastically perturbed version of the well-known Krasnoselski--Mann iteration for computing fixed points of nonexpansive maps in finite dimensional normed spaces. We discuss sufficient conditions on the stochastic noise and…

Optimization and Control · Mathematics 2023-04-04 Mario Bravo , Roberto Cominetti

Theory and application of stochastic approximation (SA) have become increasingly relevant due in part to applications in optimization and reinforcement learning. This paper takes a new look at SA with constant step-size $\alpha>0$, defined…

Statistics Theory · Mathematics 2025-11-12 Caio Kalil Lauand , Ioannis Kontoyiannis , Sean Meyn

Stochastic optimization (SO) considers the problem of optimizing an objective function in the presence of noise. Most of the solution techniques in SO estimate gradients from the noise corrupted observations of the objective and adjust…

Systems and Control · Computer Science 2018-08-03 K. Chandramouli , K. J. Prabuchandran , D. Sai Koti Reddy , Shalabh Bhatnagar

Stochastic gradient descent with momentum (SGDm) is one of the most popular optimization algorithms in deep learning. While there is a rich theory of SGDm for convex problems, the theory is considerably less developed in the context of deep…

Machine Learning · Statistics 2020-11-05 Umut Şimşekli , Lingjiong Zhu , Yee Whye Teh , Mert Gürbüzbalaban

Policy evaluation in reinforcement learning is often conducted using two-timescale stochastic approximation, which results in various gradient temporal difference methods such as GTD(0), GTD2, and TDC. Here, we provide convergence rate…

Machine Learning · Computer Science 2019-12-05 Gal Dalal , Balazs Szorenyi , Gugan Thoppe

High-dimensional linear regression under heavy-tailed noise or outlier corruption is challenging, both computationally and statistically. Convex approaches have been proven statistically optimal but suffer from high computational costs,…

Statistics Theory · Mathematics 2023-05-11 Yinan Shen , Jingyang Li , Jian-Feng Cai , Dong Xia