Related papers: Eigenvalues of the Hessian in Deep Learning: Singu…
In this paper, we study the existence and uniqueness of solutions to the weighted eigenvalue problem for $k$-Hessian equation. To achieve this, we establish the uniform a priori estimates for gradient and second derivatives of solutions to…
We propose an empirical approach centered on the spectral dynamics of weights -- the behavior of singular values and vectors during optimization -- to unify and clarify several phenomena in deep learning. We identify a consistent bias in…
The density function for the joint distribution of the first and second eigenvalues at the soft edge of unitary ensembles is found in terms of a Painlev\'e II transcendent and its associated isomonodromic system. As a corollary, the density…
Hessian captures important properties of the deep neural network loss landscape. Previous works have observed low rank structure in the Hessians of neural networks. In this paper, we propose a decoupling conjecture that decomposes the…
Many aspects of the geometry of loss functions in deep learning remain mysterious. In this paper, we work toward a better understanding of the geometry of the loss function $L$ of overparameterized feedforward neural networks. In this…
We overview some results on distributed learning with focus on a family of recently proposed algorithms known as non-Bayesian social learning. We consider different approaches to the distributed learning problem and its algorithmic…
Level curvature is a measure of sensitivity of energy levels of a disordered/chaotic system to perturbations. In the bulk of the spectrum Random Matrix Theory predicts the probability distributions of level curvatures to be given by…
We study the convergence properties of a pair of learning algorithms (learning with and without memory). This leads us to study the dominant eigenvalue of a class of random matrices. This turns out to be related to the roots of the…
The key distinguishing property of a Bayesian approach is marginalization, rather than using a single setting of weights. Bayesian marginalization can particularly improve the accuracy and calibration of modern deep neural networks, which…
While stochastic gradient descent (SGD) and variants have been surprisingly successful for training deep nets, several aspects of the optimization dynamics and generalization are still not well understood. In this paper, we present new…
In this work, we investigate the mechanism underlying loss spikes observed during neural network training. When the training enters a region with a lower-loss-as-sharper (LLAS) structure, the training becomes unstable, and the loss…
It has been observed that the statistical distribution of the eigenvalues of random matrices possesses universal properties, independent of the probability law of the stochastic matrix. In this article we find the correlation functions of…
We show that the input correlation matrix of typical classification datasets has an eigenspectrum where, after a sharp initial drop, a large number of small eigenvalues are distributed uniformly over an exponentially large range. This…
Deep learning is usually described as an experiment-driven field under continuous criticizes of lacking theoretical foundations. This problem has been partially fixed by a large volume of literature which has so far not been well organized.…
Deep Learning optimization involves minimizing a high-dimensional loss function in the weight space which is often perceived as difficult due to its inherent difficulties such as saddle points, local minima, ill-conditioning of the Hessian…
In contrast to the neatly bounded spectra of densely populated large random matrices, sparse random matrices often exhibit unbounded eigenvalue tails on the real and imaginary axis, called Lifshitz tails. In the case of asymmetric matrices,…
Over the past decades, numerous loss functions have been been proposed for a variety of supervised learning tasks, including regression, classification, ranking, and more generally structured prediction. Understanding the core principles…
The paper deals with the distribution of singular values of the input-output Jacobian of deep untrained neural networks in the limit of their infinite width. The Jacobian is the product of random matrices where the independent rectangular…
The eigenvalue density for members of the Gaussian orthogonal and unitary ensembles follows the Wigner semi-circle law. If the Gaussian entries are all shifted by a constant amount c/Sqrt(2N), where N is the size of the matrix, in the large…
Nowadays, deep learning methods, especially the convolutional neural networks (CNNs), have shown impressive performance on extracting abstract and high-level features from the hyperspectral image. However, general training process of CNNs…