Related papers: The Underlying Scaling Laws and Universal Statisti…
This work analyzes singular-value spectra of weight matrices in pretrained transformer models to understand how information is stored at both ends of the spectrum. Using Random Matrix Theory (RMT) as a zero information hypothesis, we…
We analyze statistical properties of the complex system with conditions which manifests through specific constraints on the column/row sum of the matrix elements. The presence of additional constraints besides symmetry leads to new…
It recently has been found that methods of the statistical theories of spectra can be a useful tool in the analysis of spectra far from levels of Hamiltonian systems. Several examples originate from areas, such as quantitative linguistics…
Scale-free power law structure describes complex networks derived from a wide range of real world processes. The extensive literature focuses almost exclusively on networks with power law exponent strictly larger than 2, which can be…
The population loss of trained deep neural networks often follows precise power-law scaling relations with either the size of the training dataset or the number of parameters in the network. We propose a theory that explains the origins of…
Through simple analytical calculations and numerical simulations, we demonstrate the generic existence of a self-organized macroscopic state in any large multivariate system possessing non-vanishing average correlations between a finite…
Scaling laws, a defining feature of deep learning, reveal a striking power-law improvement in model performance with increasing dataset and model size. Yet, their mathematical origins, especially the scaling exponent, have remained elusive.…
Modern Machine Learning (ML) and Deep Neural Networks (DNNs) often operate on high-dimensional data and rely on overparameterized models, where classical low-dimensional intuitions break down. In particular, the proportional regime where…
Random matrix theory (RMT) is based on two assumptions: (1) matrix-element independence, and (2) base invariance. Most of the proposed generalizations keep the first assumption and violate the second. Recently, several authors presented…
Empirically, large-scale deep learning models often satisfy a neural scaling law: the test error of the trained model improves polynomially as the model size and data size grow. However, conventional wisdom suggests the test error consists…
We consider settings where the observations are drawn from a zero-mean multivariate (real or complex) normal distribution with the population covariance matrix having eigenvalues of arbitrary multiplicity. We assume that the eigenvectors of…
Interpreting the representation and generalization powers has been a long-standing issue in the field of machine learning (ML) and artificial intelligence. This work contributes to uncovering the emergence of universal scaling laws in…
Random contractions (sub-unitary random matrices) appear naturally when considering quantized chaotic maps within a general theory of open linear stationary systems with discrete time. We analyze statistical properties of complex…
We present a random-matrix realization of a two-dimensional percolation model with the occupation probability $p$. We find that the behavior of the model is governed by the two first extreme eigenvalues. While the second extreme eigenvalue…
From benign overfitting in overparameterized models to rich power-law scalings in performance, simple ridge regression displays surprising behaviors sometimes thought to be limited to deep neural networks. This balance of phenomenological…
Scattering of electromagnetic waves in billiard-like systems has become a standard experimental tool of studying properties associated with Quantum Chaos. Random Matrix Theory (RMT) describing statistics of eigenfrequencies and associated…
Random feature maps are ubiquitous in modern statistical machine learning, where they generalize random projections by means of powerful, yet often difficult to analyze nonlinear operators. In this paper, we leverage the "concentration"…
The intrinsic dynamical complexity of classically chaotic systems enforces a universal description of the transport properties of their wave-mechanical analogues. These universal rules have been established within the framework of linear…
Complex, multivariable systems are often analyzed by grouping their constituent units into components, sometimes referred to as latent features, which afford physical or biological interpretation. However, a priori many different types of…
Random matrix theory has become a cornerstone in modern statistics and data science, providing fundamental tools for understanding high-dimensional covariance structures. Within this framework, the Wishart matrix plays a central role in…