Related papers: Full large deviation principles for the largest ei…
Stochastic gradient descent (SGD) and its variants enable modern artificial intelligence. However, theoretical understanding lags far behind their empirical success. It is widely believed that SGD has a curious ability to avoid sharp local…
We consider (annealed) large deviation principles for component empirical measures of several families of marked sparse random graphs, including (i) uniform graphs on $n$ vertices with a fixed degree distribution; (ii) uniform graphs on $n$…
We prove a large deviation principle for the expectation of macroscopic observables in quantum (and classical) Gibbs states. Our proof is based on Ruelle-Lanford functions and direct subadditivity arguments, as in the classical case,…
We consider $N\times N$ Hermitian random matrices with independent identical distributed entries. The matrix is normalized so that the average spacing between consecutive eigenvalues is of order 1/N. Under suitable assumptions on the…
Stochastic Gradient Descent (SGD) is a cornerstone of large-scale optimization, yet its theoretical behavior under heavy-tailed noise -- common in modern machine learning and reinforcement learning -- remains poorly understood. In this…
Random linear mappings are widely used in modern signal processing, compressed sensing and machine learning. These mappings may be used to embed the data into a significantly lower dimension while at the same time preserving useful…
We consider the quadratic optimization problem $$F_n^{W,h}:= \sup_{x \in S^{n-1}} ( x^T W x/2 + h^T x )\,, $$ with $W$ a (random) matrix and $h$ a random external field. We study the probabilities of large deviation of $F_n^{W,h}$ for $h$ a…
We show that the global fluctuations of spectra of GOE and GUE matrices and their principal submatrices executing Dyson's Brownian motion are Gaussian in the limit of large matrix dimensions. For nested submatrices one obtains a limiting…
We consider Hermitian random matrices of the form $H = W + \lambda V$, where $W$ is a Wigner matrix and $V$ a diagonal random matrix independent of $W$. We assume subexponential decay for the matrix entries of $W$ and we choose $\lambda…
We study the phenomenon of "crowding" near the largest eigenvalue $\lambda_{\max}$ of random $N \times N$ matrices belonging to the Gaussian Unitary Ensemble (GUE) of random matrix theory. We focus on two distinct quantities: (i) the…
We present an analytical technique to compute the probability of rare events in which the largest eigenvalue of a random matrix is atypically large (i.e.\ the right tail of its large deviations). The results also transfer to the left tail…
The probability of large deviations of the smallest Schmidt eigenvalue for random pure states of bipartite systems, denoted as $A$ and $B$, is computed analytically using a Coulomb gas method. It is shown that this probability, for large…
Given a large sample covariance matrix $S_N=\frac 1n\Gamma_N^{1/2}Z_N Z_N^*\Gamma_N^{1/2}\, ,$ where $Z_N$ is a $N\times n$ matrix with i.i.d. centered entries, and $\Gamma_N$ is a $N\times N$ deterministic Hermitian positive semidefinite…
We derive exponential bounds on probabilities of large deviations for "light tail" martingales taking values in finite-dimensional normed spaces. Our primary emphasis is on the case where the bounds are dimension-independent or nearly so.…
The $W$-random graphs provide a flexible framework for modeling large random networks. Using the Large Deviation Principle (LDP) for $W$-random graphs from [9], we prove the LDP for the corresponding class of random symmetric…
We establish a moderate deviation principle (MDP) for the number of eigenvalues of a Wigner matrix in an interval close to the edge of the spectrum. Moreover we prove a MDP for the $i$th largest eigenvalue close to the edge. The proof…
We derive the joint asymptotic distribution of the outlier eigenvalues of an additively deformed Wigner matrix $H$. Our only assumptions on the deformation are that its rank be fixed and its norm bounded. Our results extend those of [The…
Recent studies have shown that heavy tails can emerge in stochastic optimization and that the heaviness of the tails have links to the generalization error. While these studies have shed light on interesting aspects of the generalization…
We consider $N\times N$ random matrices of the form $H = W + V$ where $W$ is a real symmetric Wigner matrix and $V$ a random or deterministic, real, diagonal matrix whose entries are independent of $W$. We assume subexponential decay for…
We study wide Bayesian neural networks focusing on the rare but statistically dominant fluctuations that govern posterior concentration, beyond Gaussian-process limits. Large-deviation theory provides explicit variational objectives-rate…