中文
相关论文

相关论文: Eigenvalues of the Hessian in Deep Learning: Singu…

200 篇论文

This paper presents a comprehensive review of loss functions and performance metrics in deep learning, highlighting key developments and practical insights across diverse application areas. We begin by outlining fundamental considerations…

We characterize the phenomenon of "crowding" near the largest eigenvalue $\lambda_{\max}$ of random $N \times N$ matrices belonging to the Gaussian $\beta$-ensemble of random matrix theory, including in particular the Gaussian orthogonal…

数学物理 · 物理学 2016-01-08 Anthony Perret , Gregory Schehr

Numerous deep learning applications benefit from multi-task learning with multiple regression and classification objectives. In this paper we make the observation that the performance of such systems is strongly dependent on the relative…

计算机视觉与模式识别 · 计算机科学 2018-04-25 Alex Kendall , Yarin Gal , Roberto Cipolla

We examine the geometry of neural network training using the Jacobian of trained network parameters with respect to their initial values. Our analysis reveals low-dimensional structure in the training process which is dependent on the input…

机器学习 · 计算机科学 2024-12-12 Nora Belrose , Adam Scherlis

Self-training is a classical approach in semi-supervised learning which is successfully applied to a variety of machine learning problems. Self-training algorithm generates pseudo-labels for the unlabeled examples and progressively refines…

机器学习 · 计算机科学 2020-06-22 Samet Oymak , Talha Cihad Gulcu

We introduce a methodology for analyzing neural networks through the lens of layer-wise Hessian matrices. The local Hessian of each functional block (layer) is defined as the matrix of second derivatives of a scalar function with respect to…

机器学习 · 计算机科学 2025-11-11 Maxim Bolshim , Alexander Kugaevskikh

This paper studies the learning of linear operators between infinite-dimensional Hilbert spaces. The training data comprises pairs of random input vectors in a Hilbert space and their noisy images under an unknown self-adjoint linear…

This paper proposes a sensitivity analysis framework based on set valued mapping for deep neural networks (DNN) to understand and compute how the solutions (model weights) of DNN respond to perturbations in the training data. As a DNN may…

机器学习 · 计算机科学 2024-12-17 Xin Wang , Feilong Wang , Xuegang Ban

The success of machine learning algorithms generally depends on data representation, and we hypothesize that this is because different representations can entangle and hide more or less the different explanatory factors of variation behind…

机器学习 · 计算机科学 2014-04-24 Yoshua Bengio , Aaron Courville , Pascal Vincent

In this work, we introduce a novel probabilistic representation of deep learning, which provides an explicit explanation for the Deep Neural Networks (DNNs) in three aspects: (i) neurons define the energy of a Gibbs distribution; (ii) the…

机器学习 · 计算机科学 2019-08-27 Xinjie Lan , Kenneth E. Barner

We present the bulk-boundary decomposition as a new framework for understanding the training dynamics of deep neural networks. Starting from the stochastic gradient descent formulation, we show that the Lagrangian can be reorganized into a…

机器学习 · 计算机科学 2025-11-05 Donghee Lee , Hye-Sung Lee , Jaeok Yi

For a large class of feature maps we provide a tight asymptotic characterisation of the test error associated with learning the readout layer, in the high-dimensional limit where the input dimension, hidden layer widths, and number of…

机器学习 · 统计学 2024-06-11 Dominik Schröder , Daniil Dmitriev , Hugo Cui , Bruno Loureiro

Deep neural networks are widely used prediction algorithms whose performance often improves as the number of weights increases, leading to over-parametrization. We consider a two-layered neural network whose first layer is frozen while the…

机器学习 · 计算机科学 2023-04-10 Roman Worschech , Bernd Rosenow

Second-order methods are emerging as promising alternatives to standard first-order optimizers such as gradient descent and ADAM for training neural networks. Though the advantages of including curvature information in computing…

机器学习 · 计算机科学 2025-10-15 Conor Rowan

Sharpness-Aware Minimization (SAM) has attracted significant attention for its effectiveness in improving generalization across various tasks. However, its underlying principles remain poorly understood. In this work, we analyze SAM's…

机器学习 · 计算机科学 2025-01-23 Haocheng Luo , Tuan Truong , Tung Pham , Mehrtash Harandi , Dinh Phung , Trung Le

In a previous contribution (H.J. Stoeckmann, J. Phys. A35, 5165 (2002)), the density of states was calculated for a billiard with randomly distributed delta-like scatterers, doubly averaged over the positions of the impurities and the…

无序系统与神经网络 · 物理学 2008-11-26 Thomas Guhr , Hans-Juergen Stoeckmann

We propose an end-to-end deep learning method that learns to estimate emphysema extent from proportions of the diseased tissue. These proportions were visually estimated by experts using a standard grading system, in which grades correspond…

计算机视觉与模式识别 · 计算机科学 2018-07-24 Gerda Bortsova , Florian Dubost , Silas Ørting , Ioannis Katramados , Laurens Hogeweg , Laura Thomsen , Mathilde Wille , Marleen de Bruijne

We study some features of learning models based on "delayed" and undifferentiated reinforcement and realized by simple algorithms which may be considered of a very elementary nature. We show that a modification of the Hebb-rule works well…

凝聚态物理 · 物理学 2007-05-23 Ion-Olimpiu Stamatescu

In this article, we review the literature on statistical theories of neural networks from three perspectives: approximation, training dynamics and generative models. In the first part, results on excess risks for neural networks are…

机器学习 · 统计学 2024-09-17 Namjoon Suh , Guang Cheng

Deep Neural Networks (DNNs) rely on inherent fluctuations in their internal parameters (weights and biases) to effectively navigate the complex optimization landscape and achieve robust performance. While these fluctuations are recognized…

机器学习 · 计算机科学 2025-11-14 Darsh Pareek , Umesh Kumar , Ruthu Rao , Ravi Janjam