Related papers: Small Singular Values Matter: A Random Matrix Anal…
Representation multi-task learning (MTL) has achieved tremendous success in practice. However, the theoretical understanding of these methods is still lacking. Most existing theoretical works focus on cases where all tasks share the same…
We provide non-asymptotic, relative deviation bounds for the eigenvalues of empirical covariance and Gram matrices in general settings. Unlike typical uniform bounds, which may fail to capture the behavior of smaller eigenvalues, our…
Estimation of top singular values is one of the widely used techniques and one of the intensively researched problems in Numerical Linear Algebra and Data Science. We consider here two general questions related to this problem: How top…
The first focus of this paper is the characterization of the spectrum and the singular values of the coefficient matrix stemming from the discretization with space-time grid for a parabolic diffusion problem and from the approximation of…
The traditional method of computing singular value decomposition (SVD) of a data matrix is based on a least squares principle, thus, is very sensitive to the presence of outliers. Hence the resulting inferences across different applications…
We analyze the eigenvalues of the adjacency matrices of a wide variety of random trees. Using general, broadly applicable arguments based on the interlacing inequalities for the eigenvalues of a principal submatrix of a Hermitian matrix and…
Motivated by the problem of learning with small sample sizes, this paper shows how to incorporate into support-vector machines (SVMs) those properties that have made convolutional neural networks (CNNs) successful. Particularly important is…
The singular values $\sigma >1$ of an $n \times n$ involutory matrix $A$ appear in pairs $(\sigma, \frac{1}{\sigma}),$ while the singular values $\sigma = 1$ may appear in pairs $(1,1)$ or by themselves. The left and right singular vectors…
The spectral symbols are useful tools to analyse the eigenvalue distribution when dealing with high dimensional linear systems. Given a matrix sequence with an asymptotic symbol, the last one depends only on the spectra of the individual…
Using a nonperturbative approach we examine the large frequency asymptotics of the two-point level density correlator in weakly disordered metallic grains. This allows us to study the behavior of the two-level structure factor close to the…
There has been considerable interest in using surprisal from Transformer-based language models (LMs) as predictors of human sentence processing difficulty. Recent work has observed an inverse scaling relationship between Transformers'…
Large language models (LLMs) contain billions of parameters, yet many exact values are not essential. We show that what matters most is the relative rank of weights-whether one connection is stronger or weaker than another-rather than…
We prove an estimate on the smallest singular value of a multiplicatively and additively deformed random rectangular matrix. Suppose $n\le N \le M \le \Lambda N$ for some constant $\Lambda \ge 1$. Let $X$ be an $M\times n$ random matrix…
By singular value decomposition (SVD) of a numerically singular Hessian matrix and a numerically singular system of linear equations for the experimental data (accumulated in the respective ${\chi ^2}$ function) and constraints, least…
Large language models based on transformers have achieved great empirical successes. However, as they are deployed more widely, there is a growing need to better understand their internal mechanisms in order to make them more reliable.…
A new property, the strong singular value property, is introduced, developed, and utilized to study the problem of which lists of nonnegative real numbers occur as the singular values of a matrix with a prescribed zero-nonzero pattern.
There has been a lot of interest in understanding what information is captured by hidden representations of language models (LMs). Typically, interpretation methods i) do not guarantee that the model actually uses the encoded information,…
Random Matrix Theory (RMT) is a powerful statistical tool to model spectral fluctuations. This approach has also found fruitful application in Quantum Chromodynamics (QCD). Importantly, RMT provides very efficient means to separate…
We develop deterministic perturbation bounds for singular values and vectors of orthogonally decomposable tensors, in a spirit similar to classical results for matrices such as those due to Weyl, Davis, Kahan and Wedin. Our bounds…
Let the dimension $N$ of data and the sample size $T$ tend to $\infty$ with $N/T \to c > 0$. The spectral properties of a sample correlation matrix $\mathbf{C}$ and a sample covariance matrix $\mathbf{S}$ are asymptotically equal whenever…