中文
相关论文

相关论文: Understanding Gradient Descent through the Trainin…

200 篇论文

The Jacobian matrix (or the gradient for single-output networks) is directly related to many important properties of neural networks, such as the function landscape, stationary points, (local) Lipschitz constants and robustness to…

机器学习 · 统计学 2019-02-28 Huan Zhang , Pengchuan Zhang , Cho-Jui Hsieh

We study the map learned by a family of autoencoders trained on MNIST, and evaluated on ten different data sets created by the random selection of pixel values according to ten different distributions. Specifically, we study the eigenvalues…

机器学习 · 计算机科学 2022-01-31 Susama Agarwala , Ben Dees , Corey Lowman

Some fractals -- for instance those associated with the Mandelbrot and quadratic Julia sets -- are computed by iterating a function, and identifying the boundary between hyperparameters for which the resulting series diverges or remains…

机器学习 · 计算机科学 2024-02-12 Jascha Sohl-Dickstein

Representation learning from complex data typically involves models with a large number of parameters, which in turn require large amounts of data samples. In neural network models, model complexity grows with the number of inputs to each…

机器学习 · 计算机科学 2026-03-03 Carlos Stein Brito

Through periodic Training we can gradually buildup a reproducible responses in a disordered system where plasticity dominates over elasticity as is known in classical amorphous materials and soft matter 1, 6. Here we show that a similar…

介观与纳米尺度物理 · 物理学 2026-01-01 Madhuri Mukhopadhyay

When $k$ is a field, the classical Jacobian criterion computes the singular locus of an equidimensional, finitely generated $k$-algebra as the closed subset of an ideal generated by appropriate minors of the so-called Jacobian matrix.…

交换代数 · 数学 2024-11-06 Nawaj KC

We present two analytical formulae for estimating the sensitivity -- namely, the gradient or Jacobian -- at given realizations of an arbitrary-dimensional random vector with respect to its distributional parameters. The first formula…

机器学习 · 统计学 2025-08-14 Pi-Yueh Chuang , Ahmed Attia , Emil Constantinescu

On-line and batch learning of a perceptron in a discrete weight space, where each weight can take $2 L+1$ different values, are examined analytically and numerically. The learning algorithm is based on the training of the continuous…

统计力学 · 物理学 2009-11-07 Michal Rosen-Zvi , Ido Kanter

In suitably initialized wide networks, small learning rates transform deep neural networks (DNNs) into neural tangent kernel (NTK) machines, whose training dynamics is well-approximated by a linear weight expansion of the network at…

Recent studies have shown that many important aspects of neural network learning take place within the very earliest iterations or epochs of training. For example, sparse, trainable sub-networks emerge (Frankle et al., 2019), gradient…

机器学习 · 计算机科学 2020-02-25 Jonathan Frankle , David J. Schwab , Ari S. Morcos

Deep neural networks are highly expressive machine learning models with the ability to interpolate arbitrary datasets. Deep nets are typically optimized via first-order methods and the optimization process crucially depends on the…

机器学习 · 统计学 2019-11-12 Talha Cihad Gulcu

We present pretty detailed spectral analysis of Jacobi matrices with periodically modulated entries in the case when $0$ lies on the soft edge of the spectrum of the corresponding periodic Jacobi matrix. In particular, we show that the…

谱理论 · 数学 2018-05-09 Grzegorz Świderski

Deep neural networks can approximate functions on different types of data, from images to graphs, with varied underlying structure. This underlying structure can be viewed as the geometry of the data manifold. By extending recent advances…

机器学习 · 计算机科学 2023-01-03 Saket Tiwari , George Konidaris

Neural networks storing multiple discrete attractors are canonical models of biological memory. Previously, the dynamical stability of such networks could only be guaranteed under highly restrictive conditions. Here, we derive a theory of…

无序系统与神经网络 · 物理学 2026-01-23 Uri Cohen , Máté Lengyel

We prove that for a broad class of permutation-equivariant learning rules (including SGD, Adam, and others), the training process induces a bi-Lipschitz mapping between neurons and strongly constrains the topology of the neuron distribution…

机器学习 · 计算机科学 2025-10-06 Yongyi Yang , Tomaso Poggio , Isaac Chuang , Liu Ziyin

A quadratic approximation of neural network loss landscapes has been extensively used to study the optimization process of these networks. Though, it usually holds in a very small neighborhood of the minimum, it cannot explain many…

机器学习 · 计算机科学 2022-06-23 Chao Ma , Daniel Kunin , Lei Wu , Lexing Ying

Design of reliable systems must guarantee stability against input perturbations. In machine learning, such guarantee entails preventing overfitting and ensuring robustness of models against corruption of input data. In order to maximize…

机器学习 · 统计学 2019-08-08 Judy Hoffman , Daniel A. Roberts , Sho Yaida

We demonstrate, both analytically and numerically, that learning dynamics of neural networks is generically attracted towards a self-organized critical state. The effect can be modeled with quartic interactions between non-trainable…

统计力学 · 物理学 2021-07-09 Mikhail I. Katsnelson , Vitaly Vanchurin , Tom Westerhout

We give a rigorous analysis of the statistical behavior of gradients in a randomly initialized fully connected network N with ReLU activations. Our results show that the empirical variance of the squares of the entries in the input-output…

机器学习 · 统计学 2018-10-30 Boris Hanin

A key challenge for the machine learning community is to understand and accelerate the training dynamics of deep networks that lead to delayed generalisation and emergent robustness to input perturbations, also known as grokking. Prior work…

机器学习 · 计算机科学 2025-08-01 Thomas Walker , Ahmed Imtiaz Humayun , Randall Balestriero , Richard Baraniuk