English
Related papers

Related papers: Large deviations for the largest eigenvalue of Gau…

200 papers

We study large deviations in the context of stochastic gradient descent for one-hidden-layer neural networks with quadratic loss. We derive a quenched large deviation principle, where we condition on an initial weight measure, and an…

Probability · Mathematics 2025-01-14 Christian Hirsch , Daniel Willhalm

We study an inhomogeneous sparse random graph on [N] = {1, . . . , N } as introduced in a seminal paper by Bollobas, Janson and Riordan (2007): vertices have a type (here in a compact metric space S), and edges between different vertices…

Probability · Mathematics 2023-08-21 Luisa Andreis , Wolfgang König , Heide Langhammer , Robert I. A. Patterson

We describe the asymptotic behaviour of large degrees in random hyperbolic graphs, for all values of the curvature parameter $ \alpha$. We prove that, with high probability, the node degrees satisfy the following ordering property: the…

Probability · Mathematics 2025-03-27 Loïc Gassmann

When training neural networks with full-batch gradient descent (GD) and step size $\eta$, the largest eigenvalue of the Hessian -- the sharpness $S(\boldsymbol{\theta})$ -- rises to $2/\eta$ and hovers there, a phenomenon termed the Edge of…

Machine Learning · Computer Science 2026-04-24 Fangshuo Liao , Afroditi Kolomvaki , Anastasios Kyrillidis

Constant-stepsize stochastic approximation (SA) is widely used in learning for computational efficiency. For a fixed stepsize, the iterates typically admit a stationary distribution that is rarely tractable. Prior work shows that as the…

Machine Learning · Computer Science 2026-02-17 Zedong Wang , Yuyang Wang , Ijay Narang , Felix Wang , Yuzhou Wang , Siva Theja Maguluri

Traditional analyses of gradient descent show that when the largest eigenvalue of the Hessian, also known as the sharpness $S(\theta)$, is bounded by $2/\eta$, training is "stable" and the training loss decreases monotonically. Recent…

Machine Learning · Computer Science 2023-04-12 Alex Damian , Eshaan Nichani , Jason D. Lee

Gradient clipping is a commonly used technique to stabilize the training process of neural networks. A growing body of studies has shown that gradient clipping is a promising technique for dealing with the heavy-tailed behavior that emerged…

Machine Learning · Computer Science 2023-07-26 Shaojie Li , Yong Liu

For a large $n\times m$ Gaussian matrix, we compute the joint statistics, including large deviation tails, of generalized and total variance - the scaled log-determinant $H$ and trace $T$ of the corresponding $n\times n$ covariance matrix.…

Statistical Mechanics · Physics 2016-04-29 Fabio Deelan Cunden , Pierpaolo Vivo

This paper studies probabilistic rates of convergence for consensus+innovations type of algorithms in random, generic networks. For each node, we find a lower and also a family of upper bounds on the large deviations rate function, thus…

Information Theory · Computer Science 2022-08-11 Dragana Bajovic

We study two one-parameter families of point processes connected to random matrices: the Sine_beta and Sch_tau processes. The first one is the bulk point process limit for the Gaussian beta-ensemble. For beta=1, 2 and 4 it gives the limit…

Probability · Mathematics 2013-11-19 Diane Holcomb , Benedek Valkó

Let $N$ be the number of triangles in an Erd\H{o}s-R\'enyi graph $\mathcal{G}(n,p)$ on $n$ vertices with edge density $p=d/n,$ where $d>0$ is a fixed constant. It is well known that $N$ weakly converges to the Poisson distribution with mean…

Probability · Mathematics 2022-02-15 Shirshendu Ganguly , Ella Hiesmayr , Kyeongsik Nam

Directed last passage percolation models on the plane, where one studies the weight as well as the geometry of optimizing paths (called polymers) in a field of i.i.d. weights, are paradigm examples of models in the KPZ universality class.…

Probability · Mathematics 2017-11-01 Riddhipratim Basu , Shirshendu Ganguly , Allan Sly

The exact expression for the probability density $p_{_N}(x)$ for sums of a finite number $N$ of random independent terms is obtained. It is shown that the very tail of $p_{_N}(x)$ has a Gaussian form if and only if all the random terms are…

Probability · Mathematics 2013-05-29 Michael I. Tribelsky

Recent studies have provided both empirical and theoretical evidence illustrating that heavy tails can emerge in stochastic gradient descent (SGD) in various scenarios. Such heavy tails potentially result in iterates with diverging…

Optimization and Control · Mathematics 2021-02-23 Hongjian Wang , Mert Gürbüzbalaban , Lingjiong Zhu , Umut Şimşekli , Murat A. Erdogdu

Starting from one-point tail bounds, we establish an upper tail large deviation principle for the directed landscape at the metric level. Metrics of finite rate are in one-to-one correspondence with measures supported on a set of countably…

Probability · Mathematics 2024-05-27 Sayan Das , Duncan Dauvergne , Bálint Virág

Despite the successes of probabilistic models based on passing noise through neural networks, recent work has identified that such methods often fail to capture tail behavior accurately, unless the tails of the base distribution are…

Machine Learning · Statistics 2023-06-16 Feynman Liang , Liam Hodgkinson , Michael W. Mahoney

We investigate the privacy of {\em any} algorithm whose outputs have Gaussian distribution. This work is motivated by the prevalence of such algorithms in several useful (ML) applications, and the comparatively little research that focuses…

Cryptography and Security · Computer Science 2026-05-19 Yu Wei , Yun Lu , Malik Magdon-Ismail , Vassilis Zikas

The Pearson family of ergodic diffusions with a quadratic diffusion coefficient and a linear force are characterized by explicit dynamics of their integer moments and by explicit relaxation spectral properties towards their steady state.…

Statistical Mechanics · Physics 2023-08-14 Cecile Monthus

We prove a large deviation principle for deep neural networks with Gaussian weights and at most linearly growing activation functions, such as ReLU. This generalises earlier work, in which bounded and continuous activation functions were…

Machine Learning · Statistics 2026-02-10 Quirin Vogel

Given a graph $G$ and $p\in [0,1]$, the random subgraph $G_p$ is obtained by retaining each edge of $G$ independently with probability $p$. We show that for every $\epsilon>0$, there exists a constant $C>0$ such that the following holds.…

Combinatorics · Mathematics 2024-07-24 Sahar Diskin , Joshua Erde , Mihyun Kang , Michael Krivelevich
‹ Prev 1 4 5 6 7 8 10 Next ›