English
Related papers

Related papers: Small Singular Values Matter: A Random Matrix Anal…

200 papers

Transformers deliver outstanding performance across a wide range of tasks and are now a dominant backbone architecture for large language models (LLMs). Their task-solving performance is improved by increasing parameter size, as shown in…

Computation and Language · Computer Science 2025-06-10 Hidetaka Kamigaito , Ying Zhang , Jingun Kwon , Katsuhiko Hayashi , Manabu Okumura , Taro Watanabe

The randomized singular value decomposition (R-SVD) is a popular sketching-based algorithm for efficiently computing the partial SVD of a large matrix. When the matrix is low-rank, the R-SVD produces its partial SVD exactly; but when the…

Information Theory · Computer Science 2023-07-07 Elad Romanov

Current practices in metric evaluation focus on one single dataset, e.g., Newstest dataset in each year's WMT Metrics Shared Task. However, in this paper, we qualitatively and quantitatively show that the performances of metrics are…

Computation and Language · Computer Science 2022-04-21 Jiannan Xiang , Huayang Li , Yahui Liu , Lemao Liu , Guoping Huang , Defu Lian , Shuming Shi

We address overcrowding estimates for the singular values of random iid matrices, as well as for the eigenvalues of random Wigner matrices. We show evidence of long range separation under arbitrary perturbation even in matrices of discrete…

Probability · Mathematics 2018-10-09 Hoi H. Nguyen

Although it is known that transformer language models (LMs) pass features from early layers to later layers, it is not well understood how this information is represented and routed by the model. We analyze a mechanism used in two LMs to…

Computation and Language · Computer Science 2025-05-12 Jack Merullo , Carsten Eickhoff , Ellie Pavlick

Vector embeddings derived from large language models (LLMs) show promise in capturing latent information from the literature. Interestingly, these can be integrated into material embeddings, potentially useful for data-driven predictions of…

Computation and Language · Computer Science 2024-09-19 Luke P. J. Gilligan , Matteo Cobelli , Hasan M. Sayeed , Taylor D. Sparks , Stefano Sanvito

In this short note we collect together known results on the use of Random Matrix Theory in lattice statistical mechanics. The purpose here is two fold. Firstly the RMT analysis provides an intrinsic characterization of integrability, and…

Statistical Mechanics · Physics 2007-05-23 J. -Ch. Angles d'Auriac , J. -M. Maillard

Singular values of a data in a matrix form provide insights on the structure of the data, the effective dimensionality, and the choice of hyper-parameters on higher-level data analysis tools. However, in many practical applications such as…

Machine Learning · Statistics 2017-03-21 Ashish Khetan , Sewoong Oh

We take a first small step to extend the validity of Rudelson-Vershynin type estimates to some sparse random matrices, here random permutation matrices. We give lower (and upper) bounds on the smallest singular value of a large random…

Probability · Mathematics 2014-04-16 Gérard Ben Arous , Kim Dang

Recently, Ainsworth et al. showed that using weight matching (WM) to minimize the $L^2$ distance in a permutation search of model parameters effectively identifies permutations that satisfy linear mode connectivity (LMC), where the loss…

Machine Learning · Computer Science 2025-04-09 Akira Ito , Masanori Yamada , Atsutoshi Kumagai

In many practical situations we would like to estimate the covariance matrix of a set of variables from an insufficient amount of data. More specifically, if we have a set of $N$ independent, identically distributed measurements of an $M$…

Probability · Mathematics 2010-10-05 Thomas L. Marzetta , Gabriel H. Tucci , Steven H. Simon

The singular value decomposition (SVD) and the principal component analysis are fundamental tools and probably the most popular methods for data dimension reduction. The rapid growth in the size of data matrices has lead to a need for…

Statistics Theory · Mathematics 2020-02-03 Ting-Li Chen , Su-Yun Huang , Weichung Wang

Random matrix theory (RMT) is based on two assumptions: (1) matrix-element independence, and (2) base invariance. Most of the proposed generalizations keep the first assumption and violate the second. Recently, several authors presented…

Statistical Mechanics · Physics 2009-07-14 A. Y. Abul-Magd

We analyze how large language models (LLMs) represent out-of-context words, investigating their reliance on the given context to capture their semantics. Our likelihood-guided text perturbations reveal a correlation between token likelihood…

Computation and Language · Computer Science 2023-03-16 Valeria Ruscio , Valentino Maiorca , Fabrizio Silvestri

A recent Letter attempted to reconcile the disagreement between neutron resonance data and random matrix theory (RMT). To this end, a new formula was derived for transforming measured ({\Gamma}_{{\lambda}n}) to reduced…

Nuclear Theory · Physics 2011-01-25 P. E. Koehler , F. Bečvář , M. Krtička , J. A. Harvey , K. H. Guber

The problem of estimating the smallest singular value of random square matrices is important in connection with matrix computations and analysis of the spectral distribution. In this survey, we consider recent developments in the study of…

Probability · Mathematics 2022-06-02 Konstantin Tikhomirov

We employ the random matrix theory (RMT) framework to revisit the distribution of resonance widths in quantum chaotic systems weakly coupled to the continuum via a finite number M of open channels. In contrast to the standard first-order…

Mesoscale and Nanoscale Physics · Physics 2015-06-10 Yan V. Fyodorov , Dmitry V. Savin

It has been observed that the performances of many high-dimensional estimation problems are universal with respect to underlying sensing (or design) matrices. Specifically, matrices with markedly different constructions seem to achieve…

Information Theory · Computer Science 2023-07-24 Rishabh Dudeja , Subhabrata Sen , Yue M. Lu

Large pre-trained transformers are show-stealer in modern-day deep learning, and it becomes crucial to comprehend the parsimonious patterns that exist within them as they grow in scale. With exploding parameter counts, Lottery Ticket…

Machine Learning · Computer Science 2023-08-11 Ajay Jaiswal , Shiwei Liu , Tianlong Chen , Zhangyang Wang

Datasets from single-molecule experiments often reflect a large variety of molecular behaviour. The exploration of such datasets can be challenging, especially if knowledge about the data is limited and a priori assumptions about expected…

Data Analysis, Statistics and Probability · Physics 2020-04-06 Anton Vladyka , Tim Albrecht
‹ Prev 1 3 4 5 6 7 10 Next ›