English
Related papers

Related papers: Depth, Not Data: An Analysis of Hessian Spectral B…

200 papers

To understand the dynamics of optimization in deep neural networks, we develop a tool to study the evolution of the entire Hessian spectrum throughout the optimization process. Using this, we study a number of hypotheses concerning…

Machine Learning · Computer Science 2019-01-30 Behrooz Ghorbani , Shankar Krishnan , Ying Xiao

Given an optimization problem, the Hessian matrix and its eigenspectrum can be used in many ways, ranging from designing more efficient second-order algorithms to performing model analysis and regression diagnostics. When nonlinear models…

Machine Learning · Statistics 2021-03-18 Zhenyu Liao , Michael W. Mahoney

We study the properties of common loss surfaces through their Hessian matrix. In particular, in the context of deep learning, we empirically show that the spectrum of the Hessian is composed of two parts: (1) the bulk centered near zero,…

Machine Learning · Computer Science 2018-05-08 Levent Sagun , Utku Evci , V. Ugur Guney , Yann Dauphin , Leon Bottou

The Hessian spectrum of trained deep networks exhibits a characteristic structure: a continuous bulk of near-zero eigenvalues and a small number of large outlier eigenvalues (spikes), confirming the relevance of Random Matrix Theory in deep…

Machine Learning · Computer Science 2026-05-19 Hugo Vigna , Samuel Bontemps

We look at the eigenvalues of the Hessian of a loss function before and after training. The eigenvalue distribution is seen to be composed of two parts, the bulk which is concentrated around zero, and the edges which are scattered away from…

Machine Learning · Computer Science 2017-10-06 Levent Sagun , Leon Bottou , Yann LeCun

It is well-known that the Hessian of deep loss landscape matters to optimization, generalization, and even robustness of deep learning. Recent works empirically discovered that the Hessian spectrum in deep learning has a two-component…

Machine Learning · Computer Science 2022-08-02 Zeke Xie , Qian-Yuan Tang , Yunfeng Cai , Mingming Sun , Ping Li

We show that the input correlation matrix of typical classification datasets has an eigenspectrum where, after a sharp initial drop, a large number of small eigenvalues are distributed uniformly over an exponentially large range. This…

Machine Learning · Computer Science 2022-06-23 Rubing Yang , Jialin Mao , Pratik Chaudhari

The loss function of deep networks is known to be non-convex but the precise nature of this nonconvexity is still an active area of research. In this work, we study the loss landscape of deep networks through the eigendecompositions of…

Machine Learning · Computer Science 2019-02-08 Guillaume Alain , Nicolas Le Roux , Pierre-Antoine Manzagol

Hessian captures important properties of the deep neural network loss landscape. Previous works have observed low rank structure in the Hessians of neural networks. In this paper, we propose a decoupling conjecture that decomposes the…

Machine Learning · Computer Science 2022-10-24 Yikai Wu , Xingyu Zhu , Chenwei Wu , Annie Wang , Rong Ge

Many recent works have studied the eigenvalue spectrum of the Conjugate Kernel (CK) defined by the nonlinear feature map of a feedforward neural network. However, existing results only establish weak convergence of the empirical eigenvalue…

Machine Learning · Statistics 2024-02-16 Zhichao Wang , Denny Wu , Zhou Fan

The Hessian of a neural network captures parameter interactions through second-order derivatives of the loss. It is a fundamental object of study, closely tied to various problems in deep learning, including model design, optimization, and…

Machine Learning · Computer Science 2021-07-02 Sidak Pal Singh , Gregor Bachmann , Thomas Hofmann

Recent work has highlighted a surprising alignment between gradients and the top eigenspace of the Hessian -- termed the Dominant subspace -- during neural network training. Concurrently, there has been growing interest in the distinct…

Machine Learning · Computer Science 2025-05-20 Daniyar Zakarin , Sidak Pal Singh

The eigenvalues and eigenvectors of the connectivity matrix of complex networks contain information about its topology and its collective behavior. In particular, the spectral density $\rho(\lambda)$ of this matrix reveals important network…

Adaptation and Self-Organizing Systems · Physics 2009-11-10 M. A. M. de Aguiar , Y. Bar-Yam

Near an optimal learning point of a neural network, the learning performance of gradient descent dynamics is dictated by the Hessian matrix of the loss function with respect to the network parameters. We characterize the Hessian…

Machine Learning · Statistics 2025-12-18 Carlos Couto , José Mourão , Mário A. T. Figueiredo , Pedro Ribeiro

Neural networks (NNs) are central to modern machine learning and achieve state-of-the-art results in many applications. However, the relationship between loss geometry and generalization is still not well understood. The local geometry of…

Machine Learning · Computer Science 2026-04-15 Yuto Omae , Kazuki Sakai , Yohei Kakimoto , Makoto Sasaki , Yusuke Sakai , Hirotaka Takahashi

This study delves into the intricate dynamics of trained deep neural networks and their relationships with network parameters. Trained networks predominantly continue training in a single direction, known as the drift mode. This drift mode…

Machine Learning · Computer Science 2023-11-02 David Haink

The spectrum of the nonbacktracking matrix associated to a network is known to contain fundamental information regarding percolation properties of the network. Indeed, the inverse of its leading eigenvalue is often used as an estimate for…

Physics and Society · Physics 2025-01-30 James Martin , Tim Rogers , Luca Zanetti

Stochastic gradient descent (SGD) is widely used in deep learning due to its computational efficiency, but a complete understanding of why SGD performs so well remains a major challenge. It has been observed empirically that most…

Machine Learning · Statistics 2022-06-20 Carmina Fjellström , Kaj Nyström

We consider deep classifying neural networks. We expose a structure in the derivative of the logits with respect to the parameters of the model, which is used to explain the existence of outliers in the spectrum of the Hessian. Previous…

Machine Learning · Computer Science 2019-01-25 Vardan Papyan

We present PYHESSIAN, a new scalable framework that enables fast computation of Hessian (i.e., second-order derivative) information for deep neural networks. PYHESSIAN enables fast computations of the top Hessian eigenvalues, the Hessian…

Machine Learning · Computer Science 2021-04-21 Zhewei Yao , Amir Gholami , Kurt Keutzer , Michael Mahoney
‹ Prev 1 2 3 10 Next ›