中文
相关论文

相关论文: Understanding Gradient Descent through the Trainin…

200 篇论文

Neural models learn representations of high-dimensional data on low-dimensional manifolds. Multiple factors, including stochasticities in the training process, model architectures, and additional inductive biases, may induce different…

机器学习 · 计算机科学 2025-12-02 Hanlin Yu , Berfin Inal , Georgios Arvanitidis , Soren Hauberg , Francesco Locatello , Marco Fumero

We provide several new results on the sample complexity of vector-valued linear predictors (parameterized by a matrix), and more generally neural networks. Focusing on size-independent bounds, where only the Frobenius norm distance of the…

机器学习 · 计算机科学 2023-10-26 Roey Magen , Ohad Shamir

The paper proposes an approach to training a convolutional neural network using information on the level of distortion of input data. The learning process is modified with an additional layer, which is subsequently deleted, so the…

计算机视觉与模式识别 · 计算机科学 2019-12-03 Igor Janiszewski , Dmitry Slugin , Vladimir V. Arlazarov

We study quantum neural networks made by parametric one-qubit gates and fixed two-qubit gates in the limit of infinite width, where the generated function is the expectation value of the sum of single-qubit observables over all the qubits.…

量子物理 · 物理学 2026-05-26 Filippo Girardi , Giacomo De Palma

We propose a novel nonlinear bidirectionally coupled heterogeneous chain network whose dynamics evolve in discrete time. The backbone of the model is a pair of popular map-based neuron models, the Chialvo and the Rulkov maps. This model is…

适应与自组织系统 · 物理学 2024-05-14 Indranil Ghosh , Anjana S. Nair , Hammed Olawale Fatoyinbo , Sishu Shankar Muni

Neural network approaches that parameterize value functions have succeeded in approximating high-dimensional optimal feedback controllers when the Hamiltonian admits explicit formulas. However, many practical problems, such as the space…

最优化与控制 · 数学 2025-10-08 Eric Gelphman , Deepanshu Verma , Nicole Tianjiao Yang , Stanley Osher , Samy Wu Fung

The local geometry of high dimensional neural network loss landscapes can both challenge our cherished theoretical intuitions as well as dramatically impact the practical success of neural network training. Indeed recent works have observed…

机器学习 · 计算机科学 2019-10-15 Stanislav Fort , Surya Ganguli

Due to the rapid growth of data and computational resources, distributed optimization has become an active research area in recent years. While first-order methods seem to dominate the field, second-order methods are nevertheless attractive…

机器学习 · 计算机科学 2018-06-21 Celestine Dünner , Aurelien Lucchi , Matilde Gargiani , An Bian , Thomas Hofmann , Martin Jaggi

We study spectral properties of unbounded Jacobi matrices with periodically modulated or blended entries. Our approach is based on uniform asymptotic analysis of generalized eigenvectors. We determine when the studied operators are…

谱理论 · 数学 2022-04-08 Grzegorz Świderski , Bartosz Trojan

Based on the ideas of Optimal Control, we introduce the new basic characteristic of a bracket generating distribution, the Jacobi symbol. In contrast to the classical Tanaka symbol, the set of Jacobi symbols is discrete and classifiable. We…

微分几何 · 数学 2016-11-01 Boris Doubrov , Igor Zelenko

Most machine learning methods require tuning of hyper-parameters. For kernel ridge regression with the Gaussian kernel, the hyper-parameter is the bandwidth. The bandwidth specifies the length scale of the kernel and has to be carefully…

机器学习 · 统计学 2023-12-04 Oskar Allerbo , Rebecka Jörnsten

We consider the training process of a neural network as a dynamical system acting on the high-dimensional weight space. Each epoch is an application of the map induced by the optimization algorithm and the loss function. Using this induced…

This paper proposes Hamiltonian Learning, a novel unified framework for learning with neural networks "over time", i.e., from a possibly infinite stream of data, in an online manner, without having access to future information. Existing…

机器学习 · 计算机科学 2024-09-19 Stefano Melacci , Alessandro Betti , Michele Casoni , Tommaso Guidi , Matteo Tiezzi , Marco Gori

Recent works in deep learning have shown that integrating differentiable physics simulators into the training process can greatly improve the quality of results. Although this combination represents a more complex optimization task than…

机器学习 · 计算机科学 2022-03-22 Patrick Schnell , Philipp Holl , Nils Thuerey

Despite the popularity and success of deep learning, there is limited understanding of when, how, and why neural networks generalize to unseen examples. Since learning can be seen as extracting information from data, we formally study…

机器学习 · 计算机科学 2023-06-29 Hrayr Harutyunyan

Deep learning systems achieve remarkable empirical performance, yet the stability of the training process itself remains poorly understood. Training unfolds as a high-dimensional dynamical system in which small perturbations to…

机器学习 · 计算机科学 2026-01-21 Zhipeng Zhang , Zhenjie Yao , Kai Li , Lei Yang

We regard pre-trained residual networks (ResNets) as nonlinear systems and use linearization, a common method used in the qualitative analysis of nonlinear systems, to understand the behavior of the networks under small perturbations of the…

机器学习 · 计算机科学 2019-06-03 Kai Rothauge , Zhewei Yao , Zixi Hu , Michael W. Mahoney

A commonly used approach to study stability in a complex system is by analyzing the Jacobian matrix at an equilibrium point of a dynamical system. The equilibrium point is stable if all eigenvalues have negative real parts. Here, by…

种群与进化 · 定量生物学 2016-09-02 James P. L. Tan

Recent work has uncovered a striking phenomenon in large-capacity neural networks: they contain blocks of contiguous hidden layers with highly similar representations. This block structure has two seemingly contradictory properties: on the…

机器学习 · 计算机科学 2022-02-16 Thao Nguyen , Maithra Raghu , Simon Kornblith

Gradient descent typically converges to a single minimum of the training loss without mechanisms to explore alternative minima that may generalize better. Searching for diverse minima directly in high-dimensional parameter space is…

机器学习 · 计算机科学 2025-09-16 Akshay Vegesna , Samip Dahal