Related papers: Deterministic equivalent of the Conjugate Kernel m…
We study the eigenvalue distributions of the Conjugate Kernel and Neural Tangent Kernel associated to multi-layer feedforward neural networks. In an asymptotic regime where network width is increasing linearly in sample size, under random…
Recent work in random matrix theory (RMT) has developed the notion of deterministic equivalents: typically linear surrogate models that approximate the spectral behavior of large nonlinear random matrices, such as nonlinear feature maps in…
We investigate the fluctuations around the mean of the Stieltjes transform of the empirical spectral distribution of any selfadjoint noncommutative polynomial in a Wigner matrix and a deterministic diagonal matrix. We obtain the convergence…
For a large class of symmetric random matrices with correlated entries, selected from stationary random fields of centered and square integrable variables, we show that the limiting distribution of eigenvalue counting measure always exists…
In this article, novel deterministic equivalents for the Stieltjes transform and the Shannon transform of a class of large dimensional random matrices are provided. These results are used to characterise the ergodic rate region of multiple…
We study sample covariance matrices arising from rectangular random matrices with i.i.d. columns. It was previously known that the resolvent of these matrices admits a deterministic equivalent when the spectral parameter stays bounded away…
This paper is concerned with the asymptotic distribution of the largest eigenvalues for some nonlinear random matrix ensemble stemming from the study of neural networks. More precisely we consider $M= \frac{1}{m} YY^\top$ with $Y=f(WX)$…
Despite their immense promise in performing a variety of learning tasks, a theoretical understanding of the limitations of Deep Neural Networks (DNNs) has so far eluded practitioners. This is partly due to the inability to determine the…
A seminal work [Jacot et al., 2018] demonstrated that training a neural network under specific parameterization is equivalent to performing a particular kernel method as width goes to infinity. This equivalence opened a promising direction…
In this paper, we derive the analytical behavior of the limiting spectral distribution of non-central covariance matrices of the "general information-plus-noise" type, as studied in [14]. Through the equation defining its Stieltjes…
We establish the limiting spectral distribution of Kendall's correlation matrices in the moderate high-dimensional regime where the dimension grows slower than the sample size. Our framework allows observations to be independent but not…
We consider the spectrum of the Sample Covariance matrix $\mathbf{A}_N:= \frac{\mathbf{X}_N \mathbf{X}_N^*}{N}, $ where $\mathbf{X}_N$ is the $P\times N$ matrix with i.i.d. half-heavy tailed entries and $\frac{P}{N}\to y>0$ (the entries of…
At initialization, artificial neural networks (ANNs) are equivalent to Gaussian processes in the infinite-width limit, thus connecting them to kernel methods. We prove that the evolution of an ANN during training can also be described by a…
An equation is obtained for the Stieltjes transform of the normalized distribution of singular values of non-symmetric band random matrices in the limit when the band width and rank of the matrix simultaneously tend to infinity. Conditions…
We study an "inner-product kernel" random matrix model, whose empirical spectral distribution was shown by Xiuyuan Cheng and Amit Singer to converge to a deterministic measure in the large $n$ and $p$ limit. We provide an interpretation of…
We introduce a random matrix model where the entries are dependent across both rows and columns. More precisely, we investigate matrices of the form $\X=(X_{(i-1)n+t})_{it}\in\R^{p\times n}$ derived from a linear process $X_t=\sum_j c_j…
The Neural Tangent Kernel (NTK) is an important milestone in the ongoing effort to build a theory for deep learning. Its prediction that sufficiently wide neural networks behave as kernel methods, or equivalently as random feature models,…
An interesting approach to analyzing neural networks that has received renewed attention is to examine the equivalent kernel of the neural network. This is based on the fact that a fully connected feedforward network with one hidden layer,…
Neural Stochastic Differential Equations (NSDEs) model the drift and diffusion functions of a stochastic process as neural networks. While NSDEs are known to make accurate predictions, their uncertainty quantification properties have been…
We explore the equivalence between neural networks and kernel methods by deriving the first exact representation of any finite-size parametric classification model trained with gradient descent as a kernel machine. We compare our exact…