English
Related papers

Related papers: Understanding Gradient Descent through the Trainin…

200 papers

Neural models learn representations of high-dimensional data on low-dimensional manifolds. Multiple factors, including stochasticities in the training process, model architectures, and additional inductive biases, may induce different…

Machine Learning · Computer Science 2025-12-02 Hanlin Yu , Berfin Inal , Georgios Arvanitidis , Soren Hauberg , Francesco Locatello , Marco Fumero

We provide several new results on the sample complexity of vector-valued linear predictors (parameterized by a matrix), and more generally neural networks. Focusing on size-independent bounds, where only the Frobenius norm distance of the…

Machine Learning · Computer Science 2023-10-26 Roey Magen , Ohad Shamir

The paper proposes an approach to training a convolutional neural network using information on the level of distortion of input data. The learning process is modified with an additional layer, which is subsequently deleted, so the…

Computer Vision and Pattern Recognition · Computer Science 2019-12-03 Igor Janiszewski , Dmitry Slugin , Vladimir V. Arlazarov

We study quantum neural networks made by parametric one-qubit gates and fixed two-qubit gates in the limit of infinite width, where the generated function is the expectation value of the sum of single-qubit observables over all the qubits.…

Quantum Physics · Physics 2026-05-26 Filippo Girardi , Giacomo De Palma

We propose a novel nonlinear bidirectionally coupled heterogeneous chain network whose dynamics evolve in discrete time. The backbone of the model is a pair of popular map-based neuron models, the Chialvo and the Rulkov maps. This model is…

Adaptation and Self-Organizing Systems · Physics 2024-05-14 Indranil Ghosh , Anjana S. Nair , Hammed Olawale Fatoyinbo , Sishu Shankar Muni

Neural network approaches that parameterize value functions have succeeded in approximating high-dimensional optimal feedback controllers when the Hamiltonian admits explicit formulas. However, many practical problems, such as the space…

Optimization and Control · Mathematics 2025-10-08 Eric Gelphman , Deepanshu Verma , Nicole Tianjiao Yang , Stanley Osher , Samy Wu Fung

The local geometry of high dimensional neural network loss landscapes can both challenge our cherished theoretical intuitions as well as dramatically impact the practical success of neural network training. Indeed recent works have observed…

Machine Learning · Computer Science 2019-10-15 Stanislav Fort , Surya Ganguli

Due to the rapid growth of data and computational resources, distributed optimization has become an active research area in recent years. While first-order methods seem to dominate the field, second-order methods are nevertheless attractive…

Machine Learning · Computer Science 2018-06-21 Celestine Dünner , Aurelien Lucchi , Matilde Gargiani , An Bian , Thomas Hofmann , Martin Jaggi

We study spectral properties of unbounded Jacobi matrices with periodically modulated or blended entries. Our approach is based on uniform asymptotic analysis of generalized eigenvectors. We determine when the studied operators are…

Spectral Theory · Mathematics 2022-04-08 Grzegorz Świderski , Bartosz Trojan

Based on the ideas of Optimal Control, we introduce the new basic characteristic of a bracket generating distribution, the Jacobi symbol. In contrast to the classical Tanaka symbol, the set of Jacobi symbols is discrete and classifiable. We…

Differential Geometry · Mathematics 2016-11-01 Boris Doubrov , Igor Zelenko

Most machine learning methods require tuning of hyper-parameters. For kernel ridge regression with the Gaussian kernel, the hyper-parameter is the bandwidth. The bandwidth specifies the length scale of the kernel and has to be carefully…

Machine Learning · Statistics 2023-12-04 Oskar Allerbo , Rebecka Jörnsten

We consider the training process of a neural network as a dynamical system acting on the high-dimensional weight space. Each epoch is an application of the map induced by the optimization algorithm and the loss function. Using this induced…

Machine Learning · Computer Science 2020-06-23 Iva Manojlović , Maria Fonoberova , Ryan Mohr , Aleksandr Andrejčuk , Zlatko Drmač , Yannis Kevrekidis , Igor Mezić

This paper proposes Hamiltonian Learning, a novel unified framework for learning with neural networks "over time", i.e., from a possibly infinite stream of data, in an online manner, without having access to future information. Existing…

Machine Learning · Computer Science 2024-09-19 Stefano Melacci , Alessandro Betti , Michele Casoni , Tommaso Guidi , Matteo Tiezzi , Marco Gori

Recent works in deep learning have shown that integrating differentiable physics simulators into the training process can greatly improve the quality of results. Although this combination represents a more complex optimization task than…

Machine Learning · Computer Science 2022-03-22 Patrick Schnell , Philipp Holl , Nils Thuerey

Despite the popularity and success of deep learning, there is limited understanding of when, how, and why neural networks generalize to unseen examples. Since learning can be seen as extracting information from data, we formally study…

Machine Learning · Computer Science 2023-06-29 Hrayr Harutyunyan

Deep learning systems achieve remarkable empirical performance, yet the stability of the training process itself remains poorly understood. Training unfolds as a high-dimensional dynamical system in which small perturbations to…

Machine Learning · Computer Science 2026-01-21 Zhipeng Zhang , Zhenjie Yao , Kai Li , Lei Yang

We regard pre-trained residual networks (ResNets) as nonlinear systems and use linearization, a common method used in the qualitative analysis of nonlinear systems, to understand the behavior of the networks under small perturbations of the…

Machine Learning · Computer Science 2019-06-03 Kai Rothauge , Zhewei Yao , Zixi Hu , Michael W. Mahoney

A commonly used approach to study stability in a complex system is by analyzing the Jacobian matrix at an equilibrium point of a dynamical system. The equilibrium point is stable if all eigenvalues have negative real parts. Here, by…

Populations and Evolution · Quantitative Biology 2016-09-02 James P. L. Tan

Recent work has uncovered a striking phenomenon in large-capacity neural networks: they contain blocks of contiguous hidden layers with highly similar representations. This block structure has two seemingly contradictory properties: on the…

Machine Learning · Computer Science 2022-02-16 Thao Nguyen , Maithra Raghu , Simon Kornblith

Gradient descent typically converges to a single minimum of the training loss without mechanisms to explore alternative minima that may generalize better. Searching for diverse minima directly in high-dimensional parameter space is…

Machine Learning · Computer Science 2025-09-16 Akshay Vegesna , Samip Dahal