中文
相关论文

相关论文: Eigenvalues of the Hessian in Deep Learning: Singu…

200 篇论文

In this paper, we study the existence and uniqueness of solutions to the weighted eigenvalue problem for $k$-Hessian equation. To achieve this, we establish the uniform a priori estimates for gradient and second derivatives of solutions to…

偏微分方程分析 · 数学 2025-05-07 Rongxun He , Genggeng Huang

We propose an empirical approach centered on the spectral dynamics of weights -- the behavior of singular values and vectors during optimization -- to unify and clarify several phenomena in deep learning. We identify a consistent bias in…

The density function for the joint distribution of the first and second eigenvalues at the soft edge of unitary ensembles is found in terms of a Painlev\'e II transcendent and its associated isomonodromic system. As a corollary, the density…

经典分析与常微分方程 · 数学 2015-06-11 N. S. Witte , F. Bornemann , P. J. Forrester

Hessian captures important properties of the deep neural network loss landscape. Previous works have observed low rank structure in the Hessians of neural networks. In this paper, we propose a decoupling conjecture that decomposes the…

机器学习 · 计算机科学 2022-10-24 Yikai Wu , Xingyu Zhu , Chenwei Wu , Annie Wang , Rong Ge

Many aspects of the geometry of loss functions in deep learning remain mysterious. In this paper, we work toward a better understanding of the geometry of the loss function $L$ of overparameterized feedforward neural networks. In this…

机器学习 · 计算机科学 2020-05-19 Y. Cooper

We overview some results on distributed learning with focus on a family of recently proposed algorithms known as non-Bayesian social learning. We consider different approaches to the distributed learning problem and its algorithmic…

最优化与控制 · 数学 2016-09-27 Angelia Nedić , Alex Olshevsky , César A. Uribe

Level curvature is a measure of sensitivity of energy levels of a disordered/chaotic system to perturbations. In the bulk of the spectrum Random Matrix Theory predicts the probability distributions of level curvatures to be given by…

数学物理 · 物理学 2012-02-23 Yan V Fyodorov

We study the convergence properties of a pair of learning algorithms (learning with and without memory). This leads us to study the dominant eigenvalue of a class of random matrices. This turns out to be related to the roots of the…

概率论 · 数学 2007-05-23 Natalia Komarova , Igor Rivin

The key distinguishing property of a Bayesian approach is marginalization, rather than using a single setting of weights. Bayesian marginalization can particularly improve the accuracy and calibration of modern deep neural networks, which…

机器学习 · 计算机科学 2022-03-31 Andrew Gordon Wilson , Pavel Izmailov

While stochastic gradient descent (SGD) and variants have been surprisingly successful for training deep nets, several aspects of the optimization dynamics and generalization are still not well understood. In this paper, we present new…

机器学习 · 计算机科学 2019-07-26 Xinyan Li , Qilong Gu , Yingxue Zhou , Tiancong Chen , Arindam Banerjee

In this work, we investigate the mechanism underlying loss spikes observed during neural network training. When the training enters a region with a lower-loss-as-sharper (LLAS) structure, the training becomes unstable, and the loss…

机器学习 · 计算机科学 2024-10-08 Xiaolong Li , Zhi-Qin John Xu , Zhongwang Zhang

It has been observed that the statistical distribution of the eigenvalues of random matrices possesses universal properties, independent of the probability law of the stochastic matrix. In this article we find the correlation functions of…

凝聚态物理 · 物理学 2009-10-30 B. Eynard

We show that the input correlation matrix of typical classification datasets has an eigenspectrum where, after a sharp initial drop, a large number of small eigenvalues are distributed uniformly over an exponentially large range. This…

机器学习 · 计算机科学 2022-06-23 Rubing Yang , Jialin Mao , Pratik Chaudhari

Deep learning is usually described as an experiment-driven field under continuous criticizes of lacking theoretical foundations. This problem has been partially fixed by a large volume of literature which has so far not been well organized.…

机器学习 · 计算机科学 2021-03-12 Fengxiang He , Dacheng Tao

Deep Learning optimization involves minimizing a high-dimensional loss function in the weight space which is often perceived as difficult due to its inherent difficulties such as saddle points, local minima, ill-conditioning of the Hessian…

机器学习 · 计算机科学 2023-09-28 Rohan Kashyap

In contrast to the neatly bounded spectra of densely populated large random matrices, sparse random matrices often exhibit unbounded eigenvalue tails on the real and imaginary axis, called Lifshitz tails. In the case of asymmetric matrices,…

无序系统与神经网络 · 物理学 2025-11-07 Pietro Valigi , Joseph W. Baron , Izaak Neri , Giulio Biroli , Chiara Cammarota

Over the past decades, numerous loss functions have been been proposed for a variety of supervised learning tasks, including regression, classification, ranking, and more generally structured prediction. Understanding the core principles…

机器学习 · 统计学 2020-03-03 Mathieu Blondel , André F. T. Martins , Vlad Niculae

The paper deals with the distribution of singular values of the input-output Jacobian of deep untrained neural networks in the limit of their infinite width. The Jacobian is the product of random matrices where the independent rectangular…

机器学习 · 统计学 2022-07-13 Leonid Pastur

The eigenvalue density for members of the Gaussian orthogonal and unitary ensembles follows the Wigner semi-circle law. If the Gaussian entries are all shifted by a constant amount c/Sqrt(2N), where N is the size of the matrix, in the large…

数学物理 · 物理学 2009-04-21 Kevin E. Bassler , Peter J. Forrester , Norman E. Frankel

Nowadays, deep learning methods, especially the convolutional neural networks (CNNs), have shown impressive performance on extracting abstract and high-level features from the hyperspectral image. However, general training process of CNNs…

计算机视觉与模式识别 · 计算机科学 2020-03-12 Zhiqiang Gong , Ping Zhong , Weidong Hu