English
Related papers

Related papers: Models of Heavy-Tailed Mechanistic Universality

200 papers

Given two or more Deep Neural Networks (DNNs) with the same or similar architectures, and trained on the same dataset, but trained with different solvers, parameters, hyper-parameters, regularization, etc., can we predict which DNN will…

Machine Learning · Computer Science 2020-01-28 Charles H. Martin , Michael W. Mahoney

Unraveling the reasons behind the remarkable success and exceptional generalization capabilities of deep neural networks presents a formidable challenge. Recent insights from random matrix theory, specifically those concerning the spectral…

Machine Learning · Statistics 2023-04-10 Xuanzhe Xiao , Zeng Li , Chuanlong Xie , Fengwei Zhou

Heavy-tailed distributions are found throughout many naturally occurring phenomena. We have reviewed the models of stochastic dynamics that lead to heavy-tailed distributions (and power law distributions, in particular) including the…

Mathematical Physics · Physics 2011-05-09 Ph. Blanchard , T. Krueger , D. Volchenkov

Heavy tails are often found in practice, and yet they are an Achilles heel of a variety of mainstream random probability measures such as the Dirichlet process (DP). The first contribution of this paper focuses on characterizing the tails…

Statistics Theory · Mathematics 2025-04-04 Vianey Palacios Ramirez , Miguel de Carvalho , Luis Gutierrez Inostroza

Much research effort has been devoted to explaining the success of deep learning. Random Matrix Theory (RMT) provides an emerging way to this end: spectral analysis of large random matrices involved in a trained deep neural network (DNN)…

Machine Learning · Computer Science 2022-04-06 Xuran Meng , Jianfeng Yao

Understanding random open quantum systems is critical for characterizing the performance of large-scale quantum devices and exploring macroscopic quantum phenomena. Various features in these systems, including spectral distributions, gap…

Quantum Physics · Physics 2025-07-01 Sunkyu Yu , Xianji Piao , Namkyoo Park

Training strategies for modern deep neural networks (NNs) tend to induce a heavy-tailed (HT) empirical spectral density (ESD) in the layer weights. While previous efforts have shown that the HT phenomenon correlates with good generalization…

Machine Learning · Computer Science 2025-08-12 Vignesh Kothapalli , Tianyu Pang , Shenyang Deng , Zongmin Liu , Yaoqing Yang

Heavy-tailed metrics are common and often critical to product evaluation in the online world. While we may have samples large enough for Central Limit Theorem to kick in, experimentation is challenging due to the wide confidence interval of…

Applications · Statistics 2019-05-23 Jason , Wang , Pauline Burke

We discuss non-Gaussian random matrices whose elements are random variables with heavy-tailed probability distributions. In probability theory heavy tails of the distributions describe rare but violent events which usually have dominant…

Mathematical Physics · Physics 2009-11-08 Z. Burda , J. Jurkiewicz

Heavy-tailed distributions have been studied in statistics, random matrix theory, physics, and econometrics as models of correlated systems, among other domains. Further, heavy-tail distributed eigenvalues of the covariance matrix of the…

Machine Learning · Computer Science 2021-05-25 John Y. Shin

Deep neural networks frequently suffer from performance degradation when the training data is long-tailed because several majority classes dominate the training, resulting in a biased model. Recent studies have made a great effort in…

Computer Vision and Pattern Recognition · Computer Science 2023-05-19 Mengke Li , Yiu-ming Cheung , Juyong Jiang

Random Matrix Theory (RMT) is applied to analyze the weight matrices of Deep Neural Networks (DNNs), including both production quality, pre-trained models such as AlexNet and Inception, and smaller models trained from scratch, such as…

Machine Learning · Computer Science 2019-01-25 Charles H. Martin , Michael W. Mahoney

We analyze neural scaling laws in a solvable model of last-layer fine-tuning where targets have intrinsic, instance-heterogeneous difficulty. In our Latent Instance Difficulty (LID) model, each input's target variance is governed by a…

Machine Learning · Computer Science 2026-01-08 Noam Levi

Random matrix theory, which characterizes spectral distributions of infinitely large matrices, plays a central role across diverse fields, including high-dimensional data analysis, ecology, neuroscience, and machine learning. Among its key…

Disordered Systems and Neural Networks · Physics 2026-05-26 Arata Tomoto , Jun-nosuke Teramae

We propose a stochastic process driven by memory effect with novel distributions including both exponential and leptokurtic heavy-tailed distributions. A class of distribution is analytically derived from the continuum limit of the discrete…

Statistical Finance · Quantitative Finance 2013-05-14 Jongwook Kim , Gabjin Oh

Heckman selection model is the most popular econometric model in analysis of data with sample selection. However, selection models with Normal errors cannot accommodate heavy tails in the error distribution. Recently, Marchenko and Genton…

Computation · Statistics 2014-01-08 Peng Ding

We propose a stochastic process driven by the memory effect with novel distributions which include both exponential and leptokurtic heavy-tailed distributions. A class of the distributions is analytically derived from the continuum limit of…

Statistics Theory · Mathematics 2012-03-27 Jongwook Kim , Teppei Okumura

We characterise the learning of a mixture of two clouds of data points with generic centroids via empirical risk minimisation in the high dimensional regime, under the assumptions of generic convex loss and convex regularisation. Each cloud…

Machine Learning · Statistics 2024-03-19 Urte Adomaityte , Gabriele Sicuro , Pierpaolo Vivo

Muon has recently shown promising results in LLM training. In this work, we study how to further improve Muon. We argue that Muon's orthogonalized update rule suppresses the emergence of heavy-tailed weight spectra and over-emphasizes the…

Machine Learning · Computer Science 2026-05-25 Tianyu Pang , Yujie Fang , Zihang Liu , Shenyang Deng , Lei Hsiung , Shuhua Yu , Yaoqing Yang

We study the statistics of the largest eigenvalue lambda_max of N x N random matrices with unit variance, but power-law distributed entries, P(M_{ij})~ |M_{ij}|^{-1-mu}. When mu > 4, lambda_max converges to 2 with Tracy-Widom fluctuations…

Statistical Mechanics · Physics 2015-06-25 Giulio Biroli , Jean-Philippe Bouchaud , Marc Potters
‹ Prev 1 2 3 10 Next ›