English
Related papers

Related papers: Understanding Gradient Descent through the Trainin…

200 papers

The Jacobian matrix (or the gradient for single-output networks) is directly related to many important properties of neural networks, such as the function landscape, stationary points, (local) Lipschitz constants and robustness to…

Machine Learning · Statistics 2019-02-28 Huan Zhang , Pengchuan Zhang , Cho-Jui Hsieh

We study the map learned by a family of autoencoders trained on MNIST, and evaluated on ten different data sets created by the random selection of pixel values according to ten different distributions. Specifically, we study the eigenvalues…

Machine Learning · Computer Science 2022-01-31 Susama Agarwala , Ben Dees , Corey Lowman

Some fractals -- for instance those associated with the Mandelbrot and quadratic Julia sets -- are computed by iterating a function, and identifying the boundary between hyperparameters for which the resulting series diverges or remains…

Machine Learning · Computer Science 2024-02-12 Jascha Sohl-Dickstein

Representation learning from complex data typically involves models with a large number of parameters, which in turn require large amounts of data samples. In neural network models, model complexity grows with the number of inputs to each…

Machine Learning · Computer Science 2026-03-03 Carlos Stein Brito

Through periodic Training we can gradually buildup a reproducible responses in a disordered system where plasticity dominates over elasticity as is known in classical amorphous materials and soft matter 1, 6. Here we show that a similar…

Mesoscale and Nanoscale Physics · Physics 2026-01-01 Madhuri Mukhopadhyay

When $k$ is a field, the classical Jacobian criterion computes the singular locus of an equidimensional, finitely generated $k$-algebra as the closed subset of an ideal generated by appropriate minors of the so-called Jacobian matrix.…

Commutative Algebra · Mathematics 2024-11-06 Nawaj KC

We present two analytical formulae for estimating the sensitivity -- namely, the gradient or Jacobian -- at given realizations of an arbitrary-dimensional random vector with respect to its distributional parameters. The first formula…

Machine Learning · Statistics 2025-08-14 Pi-Yueh Chuang , Ahmed Attia , Emil Constantinescu

On-line and batch learning of a perceptron in a discrete weight space, where each weight can take $2 L+1$ different values, are examined analytically and numerically. The learning algorithm is based on the training of the continuous…

Statistical Mechanics · Physics 2009-11-07 Michal Rosen-Zvi , Ido Kanter

In suitably initialized wide networks, small learning rates transform deep neural networks (DNNs) into neural tangent kernel (NTK) machines, whose training dynamics is well-approximated by a linear weight expansion of the network at…

Machine Learning · Computer Science 2020-10-29 Stanislav Fort , Gintare Karolina Dziugaite , Mansheej Paul , Sepideh Kharaghani , Daniel M. Roy , Surya Ganguli

Recent studies have shown that many important aspects of neural network learning take place within the very earliest iterations or epochs of training. For example, sparse, trainable sub-networks emerge (Frankle et al., 2019), gradient…

Machine Learning · Computer Science 2020-02-25 Jonathan Frankle , David J. Schwab , Ari S. Morcos

Deep neural networks are highly expressive machine learning models with the ability to interpolate arbitrary datasets. Deep nets are typically optimized via first-order methods and the optimization process crucially depends on the…

Machine Learning · Statistics 2019-11-12 Talha Cihad Gulcu

We present pretty detailed spectral analysis of Jacobi matrices with periodically modulated entries in the case when $0$ lies on the soft edge of the spectrum of the corresponding periodic Jacobi matrix. In particular, we show that the…

Spectral Theory · Mathematics 2018-05-09 Grzegorz Świderski

Deep neural networks can approximate functions on different types of data, from images to graphs, with varied underlying structure. This underlying structure can be viewed as the geometry of the data manifold. By extending recent advances…

Machine Learning · Computer Science 2023-01-03 Saket Tiwari , George Konidaris

Neural networks storing multiple discrete attractors are canonical models of biological memory. Previously, the dynamical stability of such networks could only be guaranteed under highly restrictive conditions. Here, we derive a theory of…

Disordered Systems and Neural Networks · Physics 2026-01-23 Uri Cohen , Máté Lengyel

We prove that for a broad class of permutation-equivariant learning rules (including SGD, Adam, and others), the training process induces a bi-Lipschitz mapping between neurons and strongly constrains the topology of the neuron distribution…

Machine Learning · Computer Science 2025-10-06 Yongyi Yang , Tomaso Poggio , Isaac Chuang , Liu Ziyin

A quadratic approximation of neural network loss landscapes has been extensively used to study the optimization process of these networks. Though, it usually holds in a very small neighborhood of the minimum, it cannot explain many…

Machine Learning · Computer Science 2022-06-23 Chao Ma , Daniel Kunin , Lei Wu , Lexing Ying

Design of reliable systems must guarantee stability against input perturbations. In machine learning, such guarantee entails preventing overfitting and ensuring robustness of models against corruption of input data. In order to maximize…

Machine Learning · Statistics 2019-08-08 Judy Hoffman , Daniel A. Roberts , Sho Yaida

We demonstrate, both analytically and numerically, that learning dynamics of neural networks is generically attracted towards a self-organized critical state. The effect can be modeled with quartic interactions between non-trainable…

Statistical Mechanics · Physics 2021-07-09 Mikhail I. Katsnelson , Vitaly Vanchurin , Tom Westerhout

We give a rigorous analysis of the statistical behavior of gradients in a randomly initialized fully connected network N with ReLU activations. Our results show that the empirical variance of the squares of the entries in the input-output…

Machine Learning · Statistics 2018-10-30 Boris Hanin

A key challenge for the machine learning community is to understand and accelerate the training dynamics of deep networks that lead to delayed generalisation and emergent robustness to input perturbations, also known as grokking. Prior work…

Machine Learning · Computer Science 2025-08-01 Thomas Walker , Ahmed Imtiaz Humayun , Randall Balestriero , Richard Baraniuk