Related papers: Small Singular Values Matter: A Random Matrix Anal…
Transformers deliver outstanding performance across a wide range of tasks and are now a dominant backbone architecture for large language models (LLMs). Their task-solving performance is improved by increasing parameter size, as shown in…
The randomized singular value decomposition (R-SVD) is a popular sketching-based algorithm for efficiently computing the partial SVD of a large matrix. When the matrix is low-rank, the R-SVD produces its partial SVD exactly; but when the…
Current practices in metric evaluation focus on one single dataset, e.g., Newstest dataset in each year's WMT Metrics Shared Task. However, in this paper, we qualitatively and quantitatively show that the performances of metrics are…
We address overcrowding estimates for the singular values of random iid matrices, as well as for the eigenvalues of random Wigner matrices. We show evidence of long range separation under arbitrary perturbation even in matrices of discrete…
Although it is known that transformer language models (LMs) pass features from early layers to later layers, it is not well understood how this information is represented and routed by the model. We analyze a mechanism used in two LMs to…
Vector embeddings derived from large language models (LLMs) show promise in capturing latent information from the literature. Interestingly, these can be integrated into material embeddings, potentially useful for data-driven predictions of…
In this short note we collect together known results on the use of Random Matrix Theory in lattice statistical mechanics. The purpose here is two fold. Firstly the RMT analysis provides an intrinsic characterization of integrability, and…
Singular values of a data in a matrix form provide insights on the structure of the data, the effective dimensionality, and the choice of hyper-parameters on higher-level data analysis tools. However, in many practical applications such as…
We take a first small step to extend the validity of Rudelson-Vershynin type estimates to some sparse random matrices, here random permutation matrices. We give lower (and upper) bounds on the smallest singular value of a large random…
Recently, Ainsworth et al. showed that using weight matching (WM) to minimize the $L^2$ distance in a permutation search of model parameters effectively identifies permutations that satisfy linear mode connectivity (LMC), where the loss…
In many practical situations we would like to estimate the covariance matrix of a set of variables from an insufficient amount of data. More specifically, if we have a set of $N$ independent, identically distributed measurements of an $M$…
The singular value decomposition (SVD) and the principal component analysis are fundamental tools and probably the most popular methods for data dimension reduction. The rapid growth in the size of data matrices has lead to a need for…
Random matrix theory (RMT) is based on two assumptions: (1) matrix-element independence, and (2) base invariance. Most of the proposed generalizations keep the first assumption and violate the second. Recently, several authors presented…
We analyze how large language models (LLMs) represent out-of-context words, investigating their reliance on the given context to capture their semantics. Our likelihood-guided text perturbations reveal a correlation between token likelihood…
A recent Letter attempted to reconcile the disagreement between neutron resonance data and random matrix theory (RMT). To this end, a new formula was derived for transforming measured ({\Gamma}_{{\lambda}n}) to reduced…
The problem of estimating the smallest singular value of random square matrices is important in connection with matrix computations and analysis of the spectral distribution. In this survey, we consider recent developments in the study of…
We employ the random matrix theory (RMT) framework to revisit the distribution of resonance widths in quantum chaotic systems weakly coupled to the continuum via a finite number M of open channels. In contrast to the standard first-order…
It has been observed that the performances of many high-dimensional estimation problems are universal with respect to underlying sensing (or design) matrices. Specifically, matrices with markedly different constructions seem to achieve…
Large pre-trained transformers are show-stealer in modern-day deep learning, and it becomes crucial to comprehend the parsimonious patterns that exist within them as they grow in scale. With exploding parameter counts, Lottery Ticket…
Datasets from single-molecule experiments often reflect a large variety of molecular behaviour. The exploration of such datasets can be challenging, especially if knowledge about the data is limited and a priori assumptions about expected…