English
Related papers

Related papers: Muon Dynamics as a Spectral Wasserstein Flow

200 papers

We propose a custom learning algorithm for shallow over-parameterized neural networks, i.e., networks with single hidden layer having infinite width. The infinite width of the hidden layer serves as an abstraction for the…

Machine Learning · Computer Science 2023-12-19 Alexis Teter , Iman Nodozi , Abhishek Halder

We provide new convergence guarantees in Wasserstein distance for diffusion-based generative models, covering both stochastic (DDPM-like) and deterministic (DDIM-like) sampling methods. We introduce a simple framework to analyze…

Machine Learning · Computer Science 2025-11-14 Eliot Beyler , Francis Bach

We study gradient flow on the exponential loss for a classification problem with a one-layer softmax attention model, where the key and query weight matrices are trained separately. Under a separability assumption on the data, we show that…

Machine Learning · Computer Science 2024-03-14 Heejune Sheen , Siyu Chen , Tianhao Wang , Harrison H. Zhou

Spectral gradient methods, such as the Muon optimizer, modify gradient updates by preserving directional information while discarding scale, and have shown strong empirical performance in deep learning. We investigate the mechanisms…

Machine Learning · Statistics 2026-02-02 Guillaume Braun , Han Bao , Wei Huang , Masaaki Imaizumi

Machine unlearning (MU) has emerged to enhance the privacy and trustworthiness of deep neural networks. Approximate MU is a practical method for large-scale models. Our investigation into approximate MU starts with identifying the steepest…

Machine Learning · Computer Science 2024-10-01 Zhehao Huang , Xinwen Cheng , JingHao Zheng , Haoran Wang , Zhengbao He , Tao Li , Xiaolin Huang

We propose Gauss-Newton's method in function space for the solution of the Navier-Stokes equations in the physics-informed neural network (PINN) framework. Upon discretization, this yields a natural gradient method that provably mimics the…

Optimization and Control · Mathematics 2024-02-19 Anas Jnini , Flavio Vella , Marius Zeinhofer

We propose a continuous real space renormalization group transformation based on gradient flow, allowing for a numerical study of renormalization without the need for costly ensemble matching. We apply our technique in a pilot study of…

High Energy Physics - Lattice · Physics 2018-11-19 Andrea Carosso , Anna Hasenfratz , Ethan T. Neil

We propose to align distributional data from the perspective of Wasserstein means. We raise the problem of regularizing Wasserstein means and propose several terms tailored to tackle different problems. Our formulation is based on the…

Machine Learning · Computer Science 2020-02-24 Liang Mi , Wen Zhang , Yalin Wang

Commonly used $f$-divergences of measures, e.g., the Kullback-Leibler divergence, are subject to limitations regarding the support of the involved measures. A remedy is regularizing the $f$-divergence by a squared maximum mean discrepancy…

Machine Learning · Statistics 2025-04-14 Viktor Stein , Sebastian Neumayer , Nicolaj Rux , Gabriele Steidl

In this paper, we show that interventionally robust optimization problems in causal models are continuous under the $G$-causal Wasserstein distance, but may be discontinuous under the standard Wasserstein distance. This highlights the…

Machine Learning · Statistics 2025-10-20 Gabriele Visentin , Patrick Cheridito

We propose a mathematically principled PDE gradient flow framework for distributionally robust optimization (DRO). Exploiting the recent advances in the intersection of Markov Chain Monte Carlo sampling and gradient flow theory, we show…

Optimization and Control · Mathematics 2026-05-27 Zusen Xu , Jia-Jie Zhu

Despite the widespread success of Transformers across various domains, their optimization guarantees in large-scale model settings are not well-understood. This paper rigorously analyzes the convergence properties of gradient flow in…

Machine Learning · Statistics 2024-11-01 Cheng Gao , Yuan Cao , Zihao Li , Yihan He , Mengdi Wang , Han Liu , Jason Matthew Klusowski , Jianqing Fan

Normalizing flows are among the most popular paradigms in generative modeling, especially for images, primarily because we can efficiently evaluate the likelihood of a data point. This is desirable both for evaluating the fit of a model,…

Machine Learning · Computer Science 2021-06-29 Frederic Koehler , Viraj Mehta , Andrej Risteski

Wasserstein Gradient Flows (WGF) with respect to specific functionals have been widely used in the machine learning literature. Recently, neural networks have been adopted to approximate certain intractable parts of the underlying…

Machine Learning · Computer Science 2024-01-26 Huminhao Zhu , Fangyikang Wang , Chao Zhang , Hanbin Zhao , Hui Qian

In recent years, plenty of metrics have been proposed to identify networks that are free of gradient explosion and vanishing. However, due to the diversity of network components and complex serial-parallel hybrid connections in modern DNNs,…

Machine Learning · Computer Science 2020-08-13 Zhaodong Chen , Lei Deng , Bangyan Wang , Guoqi Li , Yuan Xie

The Wasserstein metric has become increasingly important in many machine learning applications such as generative modeling, image retrieval and domain adaptation. Despite its appeal, it is often too costly to compute. This has motivated…

Machine Learning · Computer Science 2025-06-04 Jonathan Bobrutsky , Amit Moscovich

The Poisson-Nernst-Planck system of equations used to model ionic transport is interpreted as a gradient flow for the Wasserstein distance and a free energy in the space of probability measures with finite second moment. A variational…

Analysis of PDEs · Mathematics 2015-09-08 David Kinderlehrer , Léonard Monsaingeon , Xiang Xu

This paper presents a parameter scan technique for BSM signal models based on normalizing flow. Normalizing flow is a type of deep learning model that transforms a simple probability distribution into a complex probability distribution as…

Data Analysis, Statistics and Probability · Physics 2024-09-23 Masahiko Saito , Masahiro Morinaga , Tomoe Kishimoto , Junichi Tanaka

We study a variant of the dynamical optimal transport problem in which the energy to be minimised is modulated by the covariance matrix of the distribution. Such transport metrics arise naturally in mean-field limits of certain ensemble…

Analysis of PDEs · Mathematics 2024-12-23 Martin Burger , Matthias Erbar , Franca Hoffmann , Daniel Matthes , André Schlichting

Wasserstein distributionally robust optimization (DRO) has recently achieved empirical success for various applications in operations research and machine learning, owing partly to its regularization effect. Although connection between…

Machine Learning · Computer Science 2020-11-02 Rui Gao , Xi Chen , Anton J. Kleywegt