English
Related papers

Related papers: Small Singular Values Matter: A Random Matrix Anal…

200 papers

Large language models (LLMs) have demonstrated remarkable capabilities, yet prohibitive parameter complexity often hinders their deployment. Existing singular value decomposition (SVD) based compression methods simply deem singular values…

Computation and Language · Computer Science 2025-02-24 Dengjie Li , Tiancheng Shen , Yao Zhou , Baisong Yang , Zhongying Liu , Masheng Yang , Bernard Ghanem , Yibo Yang , Yujie Zhong , Ming-Hsuan Yang

Large language models and deep neural networks achieve strong performance but suffer from reliability issues and high computational cost. This thesis proposes a unified framework based on spectral geometry and random matrix theory to…

Machine Learning · Computer Science 2026-01-27 Davide Ettori

Random matrix theory (RMT) successfully predicts universal statistical properties of complicated wave scattering systems in the semiclassical limit, while the random coupling model offers a complete statistical model with a simple additive…

We consider a product of an arbitrary number of independent rectangular Gaussian random matrices. We derive the mean densities of its eigenvalues and singular values in the thermodynamic limit, eventually verified numerically. These…

Statistical Mechanics · Physics 2011-06-28 Z. Burda , A. Jarosz , G. Livan , M. A. Nowak , A. Swiech

Random matrix theory allows one to deduce the eigenvalue spectrum of a large matrix given only statistical information about its elements. Such results provide insight into what factors contribute to the stability of complex dynamical…

Disordered Systems and Neural Networks · Physics 2025-01-30 Joseph W. Baron , Thomas Jun Jewell , Christopher Ryder , Tobias Galla

Random matrices have played an important role in many fields including machine learning, quantum information theory and optimization. One of the main research focuses is on the deviation inequalities for eigenvalues of random matrices.…

Probability · Mathematics 2018-10-18 Xianjie Gao , Chao Zhang , Hongwei Zhang

Modern large language models (LLMs) excel at tasks that require storing and retrieving knowledge, such as factual recall and question answering. Transformers are central to this capability because they can encode information during training…

Machine Learning · Statistics 2026-03-18 Nuri Mert Vural , Alberto Bietti , Mahdi Soltanolkotabi , Denny Wu

The Randomized Singular Value Decomposition (RSVD) is a widely used algorithm for efficiently computing low-rank approximations of large matrices, without the need to construct a full-blown SVD. Of interest, of course, is the approximation…

Numerical Analysis · Mathematics 2025-10-09 Danil Akhtiamov , Reza Ghane , Babak Hassibi

Tensor models play an increasingly prominent role in many fields, notably in machine learning. In several applications, such as community detection, topic modeling and Gaussian mixture learning, one must estimate a low-rank signal from a…

Machine Learning · Statistics 2022-06-16 José Henrique de Morais Goulart , Romain Couillet , Pierre Comon

Singular value decomposition is the key tool in the analysis and understanding of linear regularization methods. In the last decade nonlinear variational approaches such as $\ell^1$ or total variation regularizations became quite prominent…

Numerical Analysis · Mathematics 2012-11-12 Martin Benning , Martin Burger

Deep neural networks used for image classification often use convolutional filters to extract distinguishing features before passing them to a linear classifier. Most interpretability literature focuses on providing semantic meaning to…

Computer Vision and Pattern Recognition · Computer Science 2022-08-16 Brenda Praggastis , Davis Brown , Carlos Ortiz Marrero , Emilie Purvine , Madelyn Shapiro , Bei Wang

The singular value decomposition is widely used to approximate data matrices with lower rank matrices. Feng and He [Ann. Appl. Stat. 3 (2009) 1634-1654] developed tests on dimensionality of the mean structure of a data matrix based on the…

Statistics Theory · Mathematics 2014-02-28 Xingdong Feng , Xuming He

In this paper, we analyse singular values of a large $p\times n$ data matrix $\mathbf{X}_n= (\mathbf{x}_{n1},\ldots,\mathbf{x}_{nn})$ where the column $\mathbf{x}_{nj}$'s are independent $p$-dimensional vectors, possibly with different…

Statistics Theory · Mathematics 2021-08-17 Tianxing Mei , Chen Wang , Jianfeng Yao

We present a brief overview of random matrix theory (RMT) with the objectives of highlighting the computational results and applications in financial markets as complex systems. An oft-encountered problem in computational finance is the…

Statistical Finance · Quantitative Finance 2018-09-27 Hirdesh K. Pharasi , Kiran Sharma , Anirban Chakraborti , Thomas H. Seligman

Large, self-supervised transformer-based language representation models have recently received significant amounts of attention, and have produced state-of-the-art results across a variety of tasks simply by scaling up pre-training on…

Computation and Language · Computer Science 2019-10-25 Alexandre Matton , Luke de Oliveira

Random Hermitian matrices are used to model complex systems without time-reversal invariance. Adding an external source to the model can have the effect of shifting some of the matrix eigenvalues, which corresponds to shifting some of the…

Mathematical Physics · Physics 2015-05-20 Marco Bertola , Robert Buckingham , Seung-Yeop Lee , Virgil U. Pierce

Let $(\varepsilon_{t})_{t>0}$ be a sequence of independent real random vectors of $p$-dimension and let $X_T= \sum_{t=s+1}^{s+T}\varepsilon_t\varepsilon^T_{t-s}/T$ be the lag-$s$ ($s$ is a fixed positive integer) auto-covariance matrix of…

Probability · Mathematics 2018-01-23 Qinwen Wang , Jianfeng Yao

This work, based on Random Matrix Theory (RMT), introduces a novel early-stopping strategy for Transformer training dynamics. Utilizing the Power Law (PL) fit to tansformer attention matrices as a probe, we demarcate training into three…

Machine Learning · Computer Science 2025-12-30 Jing He , Hua Jiang , Cheng Li , Siqian Xin , Shuzhen Yang

We consider the eigenvalues and eigenvectors of finite, low rank perturbations of random matrices. Specifically, we prove almost sure convergence of the extreme eigenvalues and appropriate projections of the corresponding eigenvectors of…

Probability · Mathematics 2012-03-19 Florent Benaych-Georges , Raj Rao Nadakuditi

Trained attention layers exhibit striking and reproducible spectral structure of the weights, including low-rank collapse, bulk deformation, and isolated spectral outliers, yet the origin of these phenomena and their implications for…