English
Related papers

Related papers: Muon Dynamics as a Spectral Wasserstein Flow

200 papers

In this study, we use Rational-Quadratic Neural Spline Flows, a sophisticated parametrization of Normalizing Flows, for inferring posterior probability distributions in scenarios where direct evaluation of the likelihood is challenging at…

Data Analysis, Statistics and Probability · Physics 2024-01-26 Mathias El Baz , Federico Sánchez

We introduce a (de)-regularization of the Maximum Mean Discrepancy (DrMMD) and its Wasserstein gradient flow. Existing gradient flows that transport samples from source distribution to target distribution with only target samples, either…

We revisit the mean field parametrization of shallow neural networks, using signed measures on unbounded parameter spaces and duality pairings that take into account the regularity and growth of activation functions. This setting directly…

Functional Analysis · Mathematics 2025-12-17 Francesca Bartolucci , Marcello Carioni , José A. Iglesias , Yury Korolev , Emanuele Naldi , Stefano Vigogna

We introduce a model of hadronization based on invertible neural networks that faithfully reproduces a simplified version of the Lund string model for meson hadronization. Additionally, we introduce a new training method for normalizing…

High Energy Physics - Phenomenology · Physics 2024-08-14 Christian Bierlich , Phil Ilten , Tony Menzo , Stephen Mrenna , Manuel Szewc , Michael K. Wilkinson , Ahmed Youssef , Jure Zupan

The Muon optimizer has recently attracted considerable attention for its strong empirical performance and use of orthogonalized updates on matrix-shaped parameters, yet its underlying mechanisms and relationship to adaptive optimizers such…

Machine Learning · Computer Science 2026-02-05 Xianbiao Qi , Marco Chen , Jiaquan Ye , Yelin He , Rong Xiao

We prove that the sequence of marginals obtained from the iterations of the Sinkhorn algorithm or the iterative proportional fitting procedure (IPFP) on joint densities, converges to an absolutely continuous curve on the $2$-Wasserstein…

Probability · Mathematics 2026-04-21 Nabarun Deb , Young-Heon Kim , Soumik Pal , Geoffrey Schiebinger

Deep Neural Networks (DNNs) have begun to thrive in the field of automation systems, owing to the recent advancements in standardising various aspects such as architecture, optimization techniques, and regularization. In this paper, we take…

Machine Learning · Computer Science 2019-07-10 Anand Krishnamoorthy Subramanian , Nak Young Chong

We study the behavior of the Wasserstein-$2$ distance between discrete measures $\mu$ and $\nu$ in $\mathbb{R}^d$ when both measures are smoothed by small amounts of Gaussian noise. This procedure, known as Gaussian-smoothed optimal…

Statistics Theory · Mathematics 2022-06-15 Yunzi Ding , Jonathan Niles-Weed

We propose a variational finite volume scheme to approximate the solutions to Wasserstein gradient flows. The time discretization is based on an implicit linearization of the Wasserstein distance expressed thanks to Benamou-Brenier formula,…

Numerical Analysis · Mathematics 2019-07-22 Clément Cancès , Thomas O. Gallouët , Gabriele Todeschi

Muon updates matrix parameters via the matrix sign of the gradient and has shown strong empirical gains, yet its dynamics and scaling behavior remain unclear in theory. We study Muon in a linear associative memory model with softmax…

Machine Learning · Computer Science 2026-05-26 Binghui Li , Kaifei Wang , Han Zhong , Pinyan Lu , Liwei Wang

This manuscript introduces a regression-type formulation for approximating the Perron-Frobenius Operator by relying on distributional snapshots of data. These snapshots may represent densities of particles. The Wasserstein metric is…

Optimization and Control · Mathematics 2020-11-03 Amirhossein Karimi , Tryphon T. Georgiou

The conventional understanding of adversarial training in generative adversarial networks (GANs) is that the discriminator is trained to estimate a divergence, and the generator learns to minimize this divergence. We argue that despite the…

Machine Learning · Statistics 2023-08-09 Mingxuan Yi , Zhanxing Zhu , Song Liu

We propose a gradient flow procedure for generative modeling by transporting particles from an initial source distribution to a target distribution, where the gradient field on the particles is given by a noise-adaptive Wasserstein Gradient…

Machine Learning · Computer Science 2024-05-14 Alexandre Galashov , Valentin de Bortoli , Arthur Gretton

Different gradient-based methods for optimizing overparameterized models can all achieve zero training error yet converge to distinctly different solutions inducing different generalization properties. We provide the first complete…

Machine Learning · Computer Science 2025-12-08 Chen Fan , Mark Schmidt , Christos Thrampoulidis

We propose a fully discrete variational scheme for nonlinear evolution equations with gradient flow structure on the space of finite Radon measures on an interval with respect to a generalized version of the Wasserstein distance with…

Numerical Analysis · Mathematics 2016-09-29 Jonathan Zinsl , Daniel Matthes

We propose a stable method to train Wasserstein generative adversarial networks. In order to enhance stability, we consider two objective functions using the $c$-transform based on Kantorovich duality which arises in the theory of optimal…

Machine Learning · Computer Science 2021-10-28 Dohyun Kwon , Yeoneung Kim , Guido Montúfar , Insoon Yang

Physics-informed neural networks and neural operators often suffer from severe optimization difficulties caused by ill-conditioned gradients, multi-scale spectral behavior, and stiffness induced by physical constraints. Recently, the Muon…

Machine Learning · Computer Science 2026-02-19 Binghang Lu , Jiahao Zhang , Guang Lin

Covariance matrices have proven highly effective across many scientific fields. Since these matrices lie within the Symmetric Positive Definite (SPD) manifold - a Riemannian space with intrinsic non-Euclidean geometry, the primary challenge…

Machine Learning · Computer Science 2025-04-02 Rui Wang , Shaocheng Jin , Ziheng Chen , Xiaoqing Luo , Xiao-Jun Wu

We formulate well-posed continuous-time generative flows for learning distributions that are supported on low-dimensional manifolds through Wasserstein proximal regularizations of $f$-divergences. Wasserstein-1 proximal operators regularize…

Machine Learning · Statistics 2024-07-17 Hyemin Gu , Markos A. Katsoulakis , Luc Rey-Bellet , Benjamin J. Zhang

We develop a fast and scalable numerical approach to solve Wasserstein gradient flows (WGFs), particularly suitable for high-dimensional cases. Our approach is to use general reduced-order models, like deep neural networks, to parameterize…

Numerical Analysis · Mathematics 2024-05-24 Yijie Jin , Shu Liu , Hao Wu , Xiaojing Ye , Haomin Zhou