Related papers: A Random Matrix Approach to Neural Networks
Tensor models play an increasingly prominent role in many fields, notably in machine learning. In several applications, such as community detection, topic modeling and Gaussian mixture learning, one must estimate a low-rank signal from a…
We derive a differential equation that governs the evolution of the generalization gap when a deep network is trained by gradient descent. This differential equation is controlled by two quantities, a contraction factor that brings together…
We study the regularity of the score function in score-based generative models and show that it naturally adapts to the smoothness of the data distribution. Under minimal assumptions, we establish Lipschitz estimates that directly support…
Consider a random graph process where vertices are chosen from the interval $[0,1]$, and edges are chosen independently at random, but so that, for a given vertex $x$, the probability that there is an edge to a vertex $y$ decreases as the…
Let $\bm{x}_1,\cdots,\bm{x}_n$ be a random sample of size $n$ from a $p$-dimensional population distribution, where $p=p(n)\rightarrow\infty$. Consider a symmetric matrix $W=X^\top X$ with parameters $n$ and $p$, where…
Let $G$ be an $N \times N$ real matrix whose entries are independent identically distributed standard normal random variables $G_{ij} \sim \mathcal{N}(0,1)$. The eigenvalues of such matrices are known to form a two-component system…
Random Feature (RF) models are used as efficient parametric approximations of kernel methods. We investigate, by means of random matrix theory, the connection between Gaussian RF models and Kernel Ridge Regression (KRR). For a Gaussian RF…
The theory of training deep networks has become a central question of modern machine learning and has inspired many practical advancements. In particular, the gradient descent (GD) optimization algorithm has been extensively studied in…
We investigate joint spectral characteristics of a family of matrices $\mathcal F $, associated with products in the semigroup generated by $\mathcal F$. In the literature, extremal measures such as the well-known joint spectral radius and…
Gram's Law describes a pattern that frequently occurs in the distribution of the non-trivial zeros of the Riemann zeta function along the critical line. Whenever Gram's Law holds true, it reduces the difficulty of computing the…
We analyze recurrent neural networks with diagonal hidden-to-hidden weight matrices, trained with gradient descent in the supervised learning setting, and prove that gradient descent can achieve optimality \emph{without} massive…
Let $X$ be a $p\times n$ independent identically distributed real Gaussian matrix with positive mean $\mu $ and variance $\sigma^2$ entries. The goal of this paper is to investigate the largest eigenvalue of the noncentral sample covariance…
This article provides an original understanding of the behavior of a class of graph-oriented semi-supervised learning algorithms in the limit of large and numerous data. It is demonstrated that the intuition at the root of these methods…
Let X_R be the zero locus in RP^n of one or two independently and Weyl distributed random real quadratic forms (this is the same as requiring that the corresponding symmetric matrices are in the Gaussian Orthogonal Ensemble). We prove that…
Consider a random graph process with $n$ vertices corresponding to points $v_{i} \sim {Unif}[0,1]$ embedded randomly in the interval, and where edges are inserted between $v_{i}, v_{j}$ independently with probability given by the graphon…
We study the problem of learning a single neuron $\mathbf{x}\mapsto \sigma(\mathbf{w}^T\mathbf{x})$ with gradient descent (GD). All the existing positive results are limited to the case where $\sigma$ is monotonic. However, it is recently…
This paper concentrates on asymptotic properties of determinants of some random symmetric matrices. If B_{n,r} is a n x r rectangular matrix and B_{n,r}' its transpose, we study det (B_{n,r}'B_{n,r}) when n,r tends to infinity with r/n \to…
We demonstrate two new important properties of the 1-path-norm of shallow neural networks. First, despite its non-smoothness and non-convexity it allows a closed form proximal operator which can be efficiently computed, allowing the use of…
For each $N\geq 1$, let $G_N$ be a simple random graph on the set of vertices $[N]=\{1,2, ..., N\}$, which is invariant by relabeling of the vertices. The asymptotic behavior as $N$ goes to infinity of correlation functions: $$ \mathfrak…
In this paper, we explore the structure of the penultimate Gram matrix in deep neural networks, which contains the pairwise inner products of outputs corresponding to a batch of inputs. In several architectures it has been observed that…