相关论文: Analysis of One-Hidden-Layer Neural Networks via t…
This paper is concerned with the asymptotic distribution of the largest eigenvalues for some nonlinear random matrix ensemble stemming from the study of neural networks. More precisely we consider $M= \frac{1}{m} YY^\top$ with $Y=f(WX)$…
This article studies the Gram random matrix model $G=\frac1T\Sigma^{\rm T}\Sigma$, $\Sigma=\sigma(WX)$, classically found in the analysis of random feature maps and random neural networks, where $X=[x_1,\ldots,x_T]\in{\mathbb R}^{p\times…
We study norm-based uniform convergence bounds for neural networks, aiming at a tight understanding of how these are affected by the architecture and type of norm constraint, for the simple class of scalar-valued one-hidden-layer networks,…
In convolutional neural networks, the linear transformation of multi-channel two-dimensional convolutional layers with linear convolution is a block matrix with doubly Toeplitz blocks. Although a "wrapping around" operation can transform…
We analyze a simple one-hidden-layer neural network with ReLU activation functions and fixed biases, with one-dimensional input and output. We study both continuous and discrete versions of the model, and we rigorously prove the convergence…
A key property of neural networks is their capacity of adapting to data during training. Yet, our current mathematical understanding of feature learning and its relationship to generalization remain limited. In this work, we provide a…
Random tensors can be used to produce random matrices. This idea is, for instance, very natural when one studies random quantum states with the aim of exploring properties that are generically true, or true with some probability. We hereby…
We analyze recurrent neural networks with diagonal hidden-to-hidden weight matrices, trained with gradient descent in the supervised learning setting, and prove that gradient descent can achieve optimality \emph{without} massive…
For a large class of symmetric random matrices with correlated entries, selected from stationary random fields of centered and square integrable variables, we show that the limiting distribution of eigenvalue counting measure always exists…
We discuss a method of the asymptotic computation of moments of the normalized eigenvalue counting measure of random matrices of large order. The method is based on the resolvent identity and on some formulas relating expectations of…
One of the distinguishing characteristics of modern deep learning systems is that they typically employ neural network architectures that utilize enormous numbers of parameters, often in the millions and sometimes even in the billions.…
We study the asymptotic of the spectral distribution for large empirical covariance matrices composed of independent Multifractal Random Walk processes. The asymptotic is taken as the observation lag shrinks to 0. In this setting, we show…
An asymptotic technique is developed to find the Signal-to-Interference-plus-Noise-Ratio (SINR) and spectral efficiency of a link with N receiver antennas in wireless networks with non-homogeneous distributions of nodes. It is found that…
We study the Conjugate Kernel associated to a multi-layer linear-width feed-forward neural network with random weights, biases and data. We show that the empirical spectral distribution of the Conjugate Kernel converges to a deterministic…
Consider the random matrix \(\bW_n = \bB_n + n^{-1}\bX_n^*\bA_n\bX_n\), where \(\bA_n\) and \(\bB_n\) are Hermitian matrices of dimensions \(p \times p\) and \(n \times n\), respectively, and \(\bX_n\) is a \(p \times n\) random matrix with…
Among the various machine learning methods solving partial differential equations, the Random Feature Method (RFM) stands out due to its accuracy and efficiency. In this paper, we demonstrate that the approximation error of RFM exhibits…
Random Matrix Theory (RMT) is applied to analyze the weight matrices of Deep Neural Networks (DNNs), including both production quality, pre-trained models such as AlexNet and Inception, and smaller models trained from scratch, such as…
The asymptotic spectrum of graphs, introduced by Zuiddam (arXiv:1807.00169, 2018), is the space of graph parameters that are additive under disjoint union, multiplicative under the strong product, normalized and monotone under homomorphisms…
We consider sparse inhomogeneous Erd\H{o}s-R\'enyi random graph ensembles where edges are connected independently with probability $p_{ij}$. We assume that $p_{ij}= \varepsilon_N f(w_i, w_j)$ where $(w_i)_{i\ge 1}$ is a sequence of…
In this paper, we consider directly estimating the eigenvalues of precision matrix, without inverting the corresponding estimator for the eigenvalues of covariance matrix. We focus on a general asymptotic regime, i.e., the large dimensional…