Related papers: Relative concentration bounds for the spectrum of …
As context windows in large language models continue to expand, it is essential to characterize how attention behaves at extreme sequence lengths. We introduce token-sample complexity: the rate at which attention computed on $n$ tokens…
In this article, we explore the spectral properties of general random kernel matrices $[K(U_i,U_j)]_{1\leq i\neq j\leq n}$ from a Lipschitz kernel $K$ with $n$ independent random variables $U_1,U_2,\ldots, U_n$ distributed uniformly over…
Let $\{X_n\}_{n\in\N}$ be a Markov chain on a measurable space $\X$ with transition kernel $P$ and let $V:\X\r[1,+\infty)$. The Markov kernel $P$ is here considered as a linear bounded operator on the weighted-supremum space $\cB_V$…
We prove that few largest (and most important) eigenvalues of random symmetric matrices of various kinds are very strongly concentrated. This strong concentration enables us to compute the means of these eigenvalues with high precision. Our…
This paper is concerned with the asymptotic distribution of the largest eigenvalues for some nonlinear random matrix ensemble stemming from the study of neural networks. More precisely we consider $M= \frac{1}{m} YY^\top$ with $Y=f(WX)$…
Dot product kernels, such as polynomial and exponential (softmax) kernels, are among the most widely used kernels in machine learning, as they enable modeling the interactions between input features, which is crucial in applications like…
This paper systematically studies the behavior of the leading eigenvectors for independent edge undirected random graphs generated from a general latent position model whose link function is possibly infinite rank and also possibly…
In this abstract paper, we introduce a new kernel learning method by a nonparametric density estimator. The estimator consists of a group of k-centroids clusterings. Each clustering randomly selects data points with randomly selected…
For a graph representation of a dataset, a straightforward normality measure for a sample can be its graph degree. Considering a weighted graph, degree of a sample is the sum of the corresponding row's values in a similarity matrix. The…
We study the behaviors of the relative Bergman kernel metrics on holomorphic families of degenerating hyperelliptic Riemann surfaces and their Jacobian varieties. Near a node or cusp, we obtain precise asymptotic formulas with explicit…
This paper investigates the critical role of eigenalignments between the kernel matrix and learning targets in achieving robust generalization in learning problems. We establish a direct connection between generalization performance in…
We analyze gene co-expression network under the random matrix theory framework. The nearest neighbor spacing distribution of the adjacency matrix of this network follows Gaussian orthogonal statistics of random matrix theory (RMT). Spectral…
Consider a random regular graph of fixed degree $d$ with $n$ vertices. We study spectral properties of the adjacency matrix and of random Schr\"odinger operators on such a graph as $n$ tends to infinity. We prove that the integrated density…
We study the spectral properties of a class of random matrices where the matrix elements depend exponentially on the distance between uniformly and randomly distributed points. This model arises naturally in various physical contexts, such…
Recently, a new line of works has emerged to understand and improve self-attention in Transformers by treating it as a kernel machine. However, existing works apply the methods for symmetric kernels to the asymmetric self-attention,…
The entanglement spectrum, i.e., the full distribution of Schmidt eigenvalues of the reduced density matrix, contains more information than the conventional entanglement entropy and has been studied recently in several many-particle…
We examine the empirical distribution of the eigenvalues and the eigenvectors of adjacency matrices of sparse regular random graphs. We find that when the degree sequence of the graph slowly increases to infinity with the number of…
The eigendecomposition of the coupling matrix of large biological networks is central to the study of the dynamics of these networks. For neural networks, this matrix should reflect the topology of the network and conform with Dale's law…
In this work, we investigate the generalization properties of random feature methods. Our analysis extends prior results for Tikhonov regularization to a broad class of spectral regularization techniques and further generalizes the setting…
Level curvature is a measure of sensitivity of energy levels of a disordered/chaotic system to perturbations. In the bulk of the spectrum Random Matrix Theory predicts the probability distributions of level curvatures to be given by…