English
Related papers

Related papers: Eigenvalues of the Hessian in Deep Learning: Singu…

200 papers

The proper initialization of weights is crucial for the effective training and fast convergence of deep neural networks (DNNs). Prior work in this area has mostly focused on balancing the variance among weights per layer to maintain…

Machine Learning · Computer Science 2020-06-05 Maciej Skorski , Alessandro Temperoni , Martin Theobald

Given an optimization problem, the Hessian matrix and its eigenspectrum can be used in many ways, ranging from designing more efficient second-order algorithms to performing model analysis and regression diagnostics. When nonlinear models…

Machine Learning · Statistics 2021-03-18 Zhenyu Liao , Michael W. Mahoney

The loss function is arguably among the most important hyperparameters for a neural network. Many loss functions have been designed to date, making a correct choice nontrivial. However, elaborate justifications regarding the choice of the…

Machine Learning · Computer Science 2022-10-31 Simon Dräger , Jannik Dunkelau

The focus of this survey paper is on the distribution function for the largest eigenvalue in the finite N Gaussian ensembles (GOE,GUE,GSE) in the edge scaling limit of N->infinity. These limiting distribution functions are expressible in…

solv-int · Physics 2008-02-03 Craig A. Tracy , Harold Widom

The largest eigenvalue of random tensors is an important feature of systems involving disorder, equivalent to the ground state energy of glassy systems or to the injective norm of quantum states. For symmetric Gaussian random tensors of…

High Energy Physics - Theory · Physics 2024-12-16 Nicolas Delporte , Naoki Sasakura

Despite the widespread adoption of neural networks, their training dynamics remain poorly understood. We show experimentally that as the size of the dataset increases, a point forms where the magnitude of the gradient of the loss becomes…

Machine Learning · Computer Science 2024-07-23 Mark Lowell

Hessians of neural network (NN) contain essential information about the curvature of NN loss landscapes which can be used to estimate NN generalization capabilities. We have previously proposed generalization criteria that rely on the…

Machine Learning · Computer Science 2025-04-25 Nikita Gabdullin

Current methods to interpret deep learning models by generating saliency maps generally rely on two key assumptions. First, they use first-order approximations of the loss function neglecting higher-order terms such as the loss curvatures.…

Machine Learning · Computer Science 2019-06-03 Sahil Singla , Eric Wallace , Shi Feng , Soheil Feizi

In recent years, various notions of capacity and complexity have been proposed for characterizing the generalization properties of stochastic gradient descent (SGD) in deep learning. Some of the popular notions that correlate well with the…

Optimization and Control · Mathematics 2021-06-15 Mert Gurbuzbalaban , Umut Şimşekli , Lingjiong Zhu

`Double descent' delineates the generalization behaviour of models depending on the regime they belong to: under- or over-parameterized. The current theoretical understanding behind the occurrence of this phenomenon is primarily based on…

Machine Learning · Statistics 2022-03-15 Sidak Pal Singh , Aurelien Lucchi , Thomas Hofmann , Bernhard Schölkopf

Despite its empirical success, deep learning still lacks a comprehensive theoretical understanding of model fitting and generalization. This paper proposes the probability distribution (PD) learning framework to analyze the optimization and…

Machine Learning · Computer Science 2025-10-09 Binchuan Qi , Wei Gong , Li Li

In gradient descent dynamics of neural networks, the top eigenvalue of the loss Hessian (sharpness) displays a variety of robust phenomena throughout training. This includes early time regimes where the sharpness may decrease during early…

Machine Learning · Computer Science 2025-02-17 Dayal Singh Kalra , Tianyu He , Maissam Barkeshli

In the Bayesian approach to structure learning of graphical models, the equivalent sample size (ESS) in the Dirichlet prior over the model parameters was recently shown to have an important effect on the maximum-a-posteriori estimate of the…

Machine Learning · Computer Science 2012-06-18 Harald Steck

Explicit representations of the eigenvalues of the peridynamic operator have been recently derived in [5]. These representations are given in terms of generalized hypergeometric functions. Asymptotic analysis of the hypergeometric functions…

Mathematical Physics · Physics 2023-08-21 Bacim Alali , Nathan Albin , Thinh Dang

We develop a theory which describes the behaviour of eigenvalues of a class of one-dimensional random non-Hermitian operators introduced recently by Hatano and Nelson. Under general assumptions on random parameters we prove that the…

Condensed Matter · Physics 2009-10-30 Ilya Ya. Goldsheid , Boris A. Khoruzhenko

Deep neural networks have revolutionized machine learning, yet their training dynamics remain theoretically unclear-we develop a continuous-time, matrix-valued stochastic differential equation (SDE) framework that rigorously connects the…

Machine Learning · Computer Science 2026-02-10 Brian Richard Olsen , Sam Fatehmanesh , Frank Xiao , Adarsh Kumarappan , Anirudh Gajula

In this paper, we present a statistical-mechanical analysis of deep learning. We elucidate some of the essential components of deep learning---pre-training by unsupervised learning and fine tuning by supervised learning. We formulate the…

Machine Learning · Statistics 2015-06-23 Masayuki Ohzeki

We extend the method of rescaled Ward identities of Ameur-Kang-Makarov to study the distribution of eigenvalues close to a bulk singularity, i.e. a point in the interior of the droplet where the density of the classical equilibrium measure…

Mathematical Physics · Physics 2016-08-31 Yacin Ameur , Seong-Mi Seo

Loss functions are at the heart of deep learning, shaping how models learn and perform across diverse tasks. They are used to quantify the difference between predicted outputs and ground truth labels, guiding the optimization process to…

Machine Learning · Computer Science 2025-09-11 Omar Elharrouss , Yasir Mahmood , Yassine Bechqito , Mohamed Adel Serhani , Elarbi Badidi , Jamal Riffi , Hamid Tairi

Most classification models can be considered as the process of matching templates. However, when intra-class uncertainty/variability is not considered, especially for datasets containing unbalanced classes, this may lead to classification…

Computer Vision and Pattern Recognition · Computer Science 2021-04-13 He Zhu , Shan Yu