中文
相关论文

相关论文: Eigenvalues of the Hessian in Deep Learning: Singu…

200 篇论文

The local geometry of high dimensional neural network loss landscapes can both challenge our cherished theoretical intuitions as well as dramatically impact the practical success of neural network training. Indeed recent works have observed…

机器学习 · 计算机科学 2019-10-15 Stanislav Fort , Surya Ganguli

Learning in neural networks critically hinges on the intricate geometry of the loss landscape associated with a given task. Traditionally, most research has focused on finding specific weight configurations that minimize the loss. In this…

统计力学 · 物理学 2024-09-30 Margherita Mele , Roberto Menichetti , Alessandro Ingrosso , Raffaello Potestio

We introduce deep learning models to estimate the masses of the binary components of black hole mergers, $(m_1,m_2)$, and three astrophysical properties of the post-merger compact remnant, namely, the final spin, $a_f$, and the frequency…

广义相对论与量子宇宙学 · 物理学 2021-12-21 Hongyu Shen , E. A. Huerta , Eamonn O'Shea , Prayush Kumar , Zhizhen Zhao

Galaxy edges or truncations are low-surface-brightness (LSB) features located in the galaxy outskirts that delimit the distance up to where the gas density enables efficient star formation. As such, they could be interpreted as a…

星系天体物理 · 物理学 2023-12-20 Jesús Fernández , Fernando Buitrago , Benjamín Sahelices

We explore the loss landscape of fully-connected and convolutional neural networks using random, low-dimensional hyperplanes and hyperspheres. Evaluating the Hessian, $H$, of the loss function on these hypersurfaces, we observe 1) an…

机器学习 · 计算机科学 2018-11-13 Stanislav Fort , Adam Scherlis

Approximate solutions of partial differential equations (PDEs) obtained by neural networks are highly affected by hyper parameter settings. For instance, the model training strongly depends on loss function design, including the choice of…

数值分析 · 数学 2025-03-13 Hee Jun Yang , Alexander Heinlein , Hyea Hyun Kim

We suggest a loss for learning deep embeddings. The new loss does not introduce parameters that need to be tuned and results in very good embeddings across a range of datasets and problems. The loss is computed by estimating two…

计算机视觉与模式识别 · 计算机科学 2016-11-04 Evgeniya Ustinova , Victor Lempitsky

While deep learning is successful in a number of applications, it is not yet well understood theoretically. A satisfactory theoretical characterization of deep learning however, is beginning to emerge. It covers the following questions: 1)…

机器学习 · 计算机科学 2019-08-27 Tomaso Poggio , Andrzej Banburski , Qianli Liao

Full-batch gradient descent on neural networks drives the largest Hessian eigenvalue to the threshold $2/\eta$, where $\eta$ is the learning rate. This phenomenon, the Edge of Stability, has resisted a unified explanation: existing accounts…

机器学习 · 计算机科学 2026-04-23 Elon Litman

Kernel density estimation is a key component of a wide variety of algorithms in machine learning, Bayesian inference, stochastic dynamics and signal processing. However, the unsupervised density estimation technique requires tuning a…

机器学习 · 计算机科学 2025-12-17 Sunia Tanweer , Firas A. Khasawneh

A quadratic approximation of neural network loss landscapes has been extensively used to study the optimization process of these networks. Though, it usually holds in a very small neighborhood of the minimum, it cannot explain many…

机器学习 · 计算机科学 2022-06-23 Chao Ma , Daniel Kunin , Lei Wu , Lexing Ying

Hessian based measures of flatness, such as the trace, Frobenius and spectral norms, have been argued, used and shown to relate to generalisation. In this paper we demonstrate that for feed forward neural networks under the cross entropy…

机器学习 · 统计学 2020-06-17 Diego Granziol

In the context of supervised learning of a function by a neural network, we claim and empirically verify that the neural network yields better results when the distribution of the data set focuses on regions where the function to learn is…

机器学习 · 统计学 2022-09-28 Paul Novello , Gaël Poëtte , David Lugato , Pietro Congedo

The level curvature distribution function is studied both analytically and numerically for the case of T-breaking perturbations over the orthogonal ensemble. The leading correction to the shape of the curvature distribution beyond the…

介观与纳米尺度物理 · 物理学 2009-10-30 C. Basu , C. M. Canali , V. E. Kravtsov , I. V. Yurkevich

We study the value distribution and extreme values of eigenfunctions for the ``quantized cat map''. This is the quantization of a hyperbolic linear map of the torus. In a previous paper it was observed that there are quantum symmetries of…

数学物理 · 物理学 2007-05-23 Par Kurlberg , Zeev Rudnick

In this paper we argue for the fundamental importance of the value distribution: the distribution of the random return received by a reinforcement learning agent. This is in contrast to the common approach to reinforcement learning which…

机器学习 · 计算机科学 2017-07-24 Marc G. Bellemare , Will Dabney , Rémi Munos

Weight normalization (WeightNorm) is widely used in practice for the training of deep neural networks and modern deep learning libraries have built-in implementations of it. In this paper, we provide the first theoretical characterizations…

机器学习 · 计算机科学 2025-01-22 Pedro Cisneros-Velarde , Zhijie Chen , Sanmi Koyejo , Arindam Banerjee

The Sharpness Aware Minimization (SAM) optimization algorithm has been shown to control large eigenvalues of the loss Hessian and provide generalization benefits in a variety of settings. The original motivation for SAM was a modified loss…

机器学习 · 计算机科学 2023-02-20 Atish Agarwala , Yann N. Dauphin

The second-order properties of the training loss have a massive impact on the optimization dynamics of deep learning models. Fort & Scherlis (2019) discovered that a large excess of positive curvature and local convexity of the loss Hessian…

机器学习 · 计算机科学 2024-08-15 Artem Vysogorets , Anna Dawid , Julia Kempe

In many contexts, customized and weighted classification scores are designed in order to evaluate the goodness of the predictions carried out by neural networks. However, there exists a discrepancy between the maximization of such scores…

机器学习 · 计算机科学 2023-05-24 Francesco Marchetti , Sabrina Guastavino , Cristina Campi , Federico Benvenuto , Michele Piana