中文
相关论文

相关论文: The Riemannian Geometry Associated to Gradient Flo…

200 篇论文

We study the convergence of gradient flows related to learning deep linear neural networks (where the activation function is the identity map) from data. In this case, the composition of the network layers amounts to simply multiplying the…

最优化与控制 · 数学 2020-10-16 Bubacarr Bah , Holger Rauhut , Ulrich Terstiege , Michael Westdickenberg

This article provides an expository account of training dynamics in the Deep Linear Network (DLN) from the perspective of the geometric theory of dynamical systems. Rigorous results by several authors are unified into a thermodynamic…

神经与进化计算 · 计算机科学 2024-11-15 Govind Menon

We derive the system of differential equations for the gradient flow characterizing the training process of linear in-context learning in full generality. Next, we explore the geometric structure of the gradient flows in two instances,…

动力系统 · 数学 2024-12-24 Songtao Lu , Yingdong Lu , Tomasz Nowicki

Non-Euclidean constraints are inherent in many kinds of data in computer vision and machine learning, typically as a result of specific invariance requirements that need to be respected during high-level inference. Often, these geometric…

计算机视觉与模式识别 · 计算机科学 2017-09-26 Suhas Lohit , Pavan Turaga

We consider the scenario of supervised learning in Deep Learning (DL) networks, and exploit the arbitrariness of choice in the Riemannian metric relative to which the gradient descent flow can be defined (a general fact of differential…

机器学习 · 计算机科学 2026-05-26 Thomas Chen

The deep linear network (DLN) is a model for implicit regularization in gradient based optimization of overparametrized learning architectures. Training the DLN corresponds to a Riemannian gradient flow, where the Riemannian metric is…

动力系统 · 数学 2023-05-12 Nadav Cohen , Govind Menon , Zsolt Veraszto

Gradient-based methods successfully train highly overparameterized models in practice, even though the associated optimization problems are markedly nonconvex. Understanding the mechanisms that make such methods effective has become a…

机器学习 · 计算机科学 2026-01-21 Hippolyte Labarrière , Cesare Molinari , Lorenzo Rosasco , Cristian Vega , Silvia Villa

We study the Riemannian geometry of the Deep Linear Network (DLN) as a foundation for a thermodynamic description of the learning process. The main tools are the use of group actions to analyze overparametrization and the use of Riemannian…

机器学习 · 计算机科学 2026-05-22 Govind Menon , Tianmin Yu

Graphs are ubiquitous, and learning on graphs has become a cornerstone in artificial intelligence and data mining communities. Unlike pixel grids in images or sequential structures in language, graphs exhibit a typical non-Euclidean…

机器学习 · 计算机科学 2026-02-12 Li Sun , Qiqi Wan , Suyang Zhou , Zhenhao Huang , Philip S. Yu

Neural networks are playing a crucial role in everyday life, with the most modern generative models able to achieve impressive results. Nonetheless, their functioning is still not very clear, and several strategies have been adopted to…

微分几何 · 数学 2024-04-10 Alessandro Benfenati , Alessio Marta

Many tasks require mapping continuous input data (e.g. images) to discrete task outputs (e.g. class labels). Yet, how neural networks learn to perform such discrete computations on continuous data manifolds remains poorly understood. Here,…

机器学习 · 计算机科学 2025-12-02 Julian Brandon , Angus Chadwick , Arthur Pellegrino

Assuming a-priori a smooth generating vector field, we introduce a generally covariant measure of the flow geometry called the referential gradient of the flow. The main result is the explicit relation between the referential gradient and…

数学物理 · 物理学 2014-11-21 J. K. Edmondson

We study the convergence of gradient flow for the training of deep neural networks. If Residual Neural Networks are a popular example of very deep architectures, their training constitutes a challenging optimization problem due notably to…

机器学习 · 计算机科学 2025-07-22 Raphaël Barboni , Gabriel Peyré , François-Xavier Vialard

Neural networks trained via gradient descent with random initialization and without any regularization enjoy good generalization performance in practice despite being highly overparametrized. A promising direction to explain this phenomenon…

机器学习 · 计算机科学 2022-05-17 Hancheng Min , Salma Tarmoun , Rene Vidal , Enrique Mallada

A complete understanding of the widely used over-parameterized deep networks is a key step for AI. In this work we try to give a geometric picture of over-parameterized deep networks using our geometrization scheme. We show that the…

机器学习 · 计算机科学 2019-02-12 Xiao Dong , Ling Zhou

The gradient flow is the evolution of fields and physical quantities along a dimensionful parameter~$t$, the flow time. We give a simple argument that relates this gradient flow and the Wilsonian renormalization group (RG) flow. We then…

高能物理 - 理论 · 物理学 2021-07-09 Hiroki Makino , Okuto Morikawa , Hiroshi Suzuki

Convolutional neural networks are widely used in imaging and image recognition. Learning such networks from training data leads to the minimization of a non-convex function. This makes the analysis of standard optimization methods such as…

最优化与控制 · 数学 2026-01-14 Jona-Maria Diederen , Holger Rauhut , Ulrich Terstiege

The paper addresses the problem of learning a regression model parameterized by a fixed-rank positive semidefinite matrix. The focus is on the nonlinear nature of the search space and on scalability to high-dimensional problems. The…

机器学习 · 计算机科学 2011-02-01 Gilles Meyer , Silvere Bonnabel , Rodolphe Sepulchre

A fundamental challenge in the theory of deep learning is to understand whether gradient-based training can promote parameters belonging to certain lower-dimensional structures (e.g., sparse or low-rank sets), leading to so-called implicit…

机器学习 · 计算机科学 2026-03-16 Sibylle Marcotte , Gabriel Peyré , Rémi Gribonval

Real world data often lie on low-dimensional Riemannian manifolds embedded in high-dimensional spaces. This motivates learning degenerate normalizing flows that map between the ambient space and a low-dimensional latent space. However, if…

机器学习 · 计算机科学 2026-04-14 Hanlin Yu , Søren Hauberg , Marcelo Hartmann , Arto Klami , Georgios Arvanitidis
‹ 上一页 1 2 3 10 下一页 ›