Related papers: A Random Matrix Approach to Neural Networks
Consider the random matrix $\Sigma = D^{1/2} X \widetilde D^{1/2}$ where $D$ and $\widetilde D$ are deterministic Hermitian nonnegative matrices with respective dimensions $N \times N$ and $n \times n$, and where $X$ is a random matrix with…
We are concerned with an approximation problem for a symmetric positive semidefinite matrix due to motivation from a class of nonlinear machine learning methods. We discuss an approximation approach that we call {matrix ridge…
Attention based neural networks are state of the art in a large range of applications. However, their performance tends to degrade when the number of layers increases. In this work, we show that enforcing Lipschitz continuity by normalizing…
This work investigates the asymptotic behaviour of the gradient approximation method called the generalized simplex gradient (GSG). This method has an error bound that at first glance seems to tend to infinity as the number of sample points…
Random Matrix Theory (RMT) is applied to analyze weight matrices of Deep Neural Networks (DNNs), including both production quality, pre-trained models such as AlexNet and Inception, and smaller models trained from scratch, such as LeNet5…
We observe a sample of $n$ independent $p$-dimensional Gaussian vectors with Toeplitz covariance matrix $ \Sigma = [\sigma_{|i-j|}]_{1 \leq i,j \leq p}$ and $\sigma_0=1$. We consider the problem of testing the hypothesis that $\Sigma$ is…
We consider properties of determinants of some random symmetric matrices issued from multivariate statistics: Wishart/Laguerre ensemble (sample covariance matrices), Uniform Gram ensemble (sample correlation matrices) and Jacobi ensemble…
We first study the generalization error of models that use a fixed feature representation (frozen intermediate layers) followed by a trainable readout layer. This setting encompasses a range of architectures, from deep random-feature models…
We study the normalized eigenvalue counting measure d\sigma of matrices of long-range percolation model. These are (2n+1)\times (2n+1) random real symmetric matrices H=\{H(i,j)\}_{i,j} whose elements are independent random variables taking…
By the continuous mapping theorem, if a sequence of $d$-dimensional random vectors $(\mathbf{W}_n)_{n\geq1}$ converges in distribution to a multivariate normal random variable $\Sigma^{1/2}\mathbf{Z}$, then the sequence of random variables…
We study the task of learning Generalized Linear models (GLMs) in the agnostic model under the Gaussian distribution. We give the first polynomial-time algorithm that achieves a constant-factor approximation for \textit{any} monotone…
In this paper, we study the equilibrium states of a $N\times N$ stochastic complex random matrix $M$, whose entries evolve in time accordingly with a Langevin equation including both Gaussian white noises and a linear disorder, materialized…
We discuss a method of the asymptotic computation of moments of the normalized eigenvalue counting measure of random matrices of large order. The method is based on the resolvent identity and on some formulas relating expectations of…
In this paper, we investigate a two-layer fully connected neural network of the form $f(X)=\frac{1}{\sqrt{d_1}}\boldsymbol{a}^\top \sigma\left(WX\right)$, where $X\in\mathbb{R}^{d_0\times n}$ is a deterministic data matrix,…
The asymptotic behaviour of Linear Spectral Statistics (LSS) of the smoothed periodogram estimator of the spectral coherency matrix of a complex Gaussian high-dimensional time series $(\y_n)_{n \in \mathbb{Z}}$ with independent components…
We consider $N\times N$ Gaussian random matrices, whose average density of eigenvalues has the Wigner semi-circle form over $[-\sqrt{2},\sqrt{2}]$. For such matrices, using a Coulomb gas technique, we compute the large $N$ behavior of the…
Random Matrix Theory (RMT) is applied to analyze the weight matrices of Deep Neural Networks (DNNs), including both production quality, pre-trained models such as AlexNet and Inception, and smaller models trained from scratch, such as…
We investigate the effect of explicitly enforcing the Lipschitz continuity of neural networks with respect to their inputs. To this end, we provide a simple technique for computing an upper bound to the Lipschitz constant---for multiple…
We consider sample covariance matrices $S_N=\frac{1}{p}\Sigma_N^{1/2}X_NX_N^* \Sigma_N^{1/2}$ where $X_N$ is a $N \times p$ real or complex matrix with i.i.d. entries with finite $12^{\rm th}$ moment and $\Sigma_N$ is a $N \times N$…
Let $G$ be a compact Lie group, $N\geq 1$ and $L>0$. The random geometric graph on $G$ is the random graph $\Gamma(N,L)$ whose vertices are $N$ random points $g_1,\ldots,g_N$ chosen under the Haar measure of $G$, and whose edges are the…