English
Related papers

Related papers: (De)-regularized Maximum Mean Discrepancy Gradient…

200 papers

Multi-modal distributions are commonly used to model clustered data in statistical learning tasks. In this paper, we consider the Mixed Linear Regression (MLR) problem. We propose an optimal transport-based framework for MLR problems,…

Machine Learning · Statistics 2021-06-17 Theo Diamandis , Yonina C. Eldar , Alireza Fallah , Farzan Farnia , Asuman Ozdaglar

We develop a geometric convergence theory for neural-network optimization within the minimizing movement scheme (MMS) framework. Reformulating each neural MMS step as a minimization over the set of increments in a Hilbert space, we show…

Optimization and Control · Mathematics 2026-05-28 Shixin Zheng , Yiwei Wang , Haizhao Yang

Gradient normalization stabilizes deep-learning optimization, and spectral normalizations are especially natural for matrix-shaped parameter blocks; Muon is the motivating example. We study an idealized deterministic, continuous-time,…

Optimization and Control · Mathematics 2026-05-11 Gabriel Peyré

Wasserstein gradient flows have become a central tool for optimization problems over probability measures. A natural numerical approach is forward-Euler time discretization. We show, however, that even in the simple case where the energy…

Numerical Analysis · Mathematics 2025-10-16 Yewei Xu , Qin Li

We develop a quantitative statistical theory of transformers in the large-context regime by adopting the abstraction of contextual flow maps (CFMs): dynamical systems that evolve a distinguished token in the presence of a contextual measure…

Machine Learning · Computer Science 2026-05-19 Shi Chen , Zhengjiang Lin , Kaizhao Liu , Philippe Rigollet

We develop a fast and scalable numerical approach to solve Wasserstein gradient flows (WGFs), particularly suitable for high-dimensional cases. Our approach is to use general reduced-order models, like deep neural networks, to parameterize…

Numerical Analysis · Mathematics 2024-05-24 Yijie Jin , Shu Liu , Hao Wu , Xiaojing Ye , Haomin Zhou

We study the discretization of generalized Wasserstein distances with nonlinear mobilities on the real line via suitable discrete metrics on the cone of N ordered particles, a setting which naturally appears in the framework of…

Analysis of PDEs · Mathematics 2022-09-01 Simone Di Marino , Lorenzo Portinale , Emanuela Radici

We propose a distributed nonparametric algorithm for solving measure-valued optimization problems with additive objectives. Such problems arise in several contexts in stochastic learning and control including Langevin sampling from an…

Optimization and Control · Mathematics 2022-02-21 Iman Nodozi , Abhishek Halder

Recently, the rectified flow (RF) has emerged as the new state-of-the-art among flow-based diffusion models due to its high efficiency advantage in straight path sampling, especially with the amazing images generated by a series of RF…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Zhiyuan Ma , Ruixun Liu , Sixian Liu , Jianjun Li , Bowen Zhou

The paper introduces a new kernel-based Maximum Mean Discrepancy (MMD) statistic for measuring the distance between two distributions given finitely-many multivariate samples. When the distributions are locally low-dimensional, the proposed…

Machine Learning · Statistics 2018-09-03 Xiuyuan Cheng , Alexander Cloninger , Ronald R. Coifman

We examine the infinite-dimensional optimization problem of finding a decomposition of a probability measure into K probability sub-measures to minimize specific loss functions inspired by applications in clustering and user grouping. We…

Optimization and Control · Mathematics 2024-06-04 Jiangze Han , Christopher Thomas Ryan , Xin T. Tong

Discrete flow models offer a powerful framework for learning distributions over discrete state spaces and have demonstrated superior performance compared to the discrete diffusion models. However, their convergence properties and error…

Statistics Theory · Mathematics 2026-05-27 Zhengyan Wan , Yidong Ouyang , Qiang Yao , Liyan Xie , Fang Fang , Hongyuan Zha , Guang Cheng

We study the global convergence of policy gradient for infinite-horizon, continuous state and action space, and entropy-regularized Markov decision processes (MDPs). We consider a softmax policy with (one-hidden layer) neural network…

Optimization and Control · Mathematics 2022-06-17 Bekzhan Kerimkulov , James-Michael Leahy , David Šiška , Lukasz Szpruch

Normalizing flows model a complex target distribution in terms of a bijective transform operating on a simple base distribution. As such, they enable tractable computation of a number of important statistical quantities, particularly…

Machine Learning · Computer Science 2022-09-01 Chandramouli Shama Sastry , Andreas Lehrmann , Marcus Brubaker , Alexander Radovic

Modern data analyses frequently encounter settings where samples of variables are contaminated by measurement error. Ignoring measurement noise can substantially degrade statistical inference, while existing correction techniques are often…

Methodology · Statistics 2026-04-15 Ritwik Vashistha , Jeff M. Phillips , Abhra Sarkar , Arya Farahi

Solving Fredholm equations of the first kind is crucial in many areas of the applied sciences. In this work we adopt a probabilistic and variational point of view by considering a minimization problem in the space of probability measures…

Optimization and Control · Mathematics 2024-05-17 Francesca R. Crucinio , Valentin De Bortoli , Arnaud Doucet , Adam M. Johansen

Decentralized stochastic optimization has emerged as a fundamental paradigm for large-scale machine learning. However, practical implementations often rely on biased gradient estimators arising from communication compression or inexact…

Optimization and Control · Mathematics 2026-04-10 Qing Xu , Yiwei Liao , Wenqi Fan , Xingxing You , Songyi Dian

Wasserstein gradient flow has emerged as a promising approach to solve optimization problems over the space of probability distributions. A recent trend is to use the well-known JKO scheme in combination with input convex neural networks to…

Machine Learning · Computer Science 2022-07-26 Jiaojiao Fan , Qinsheng Zhang , Amirhossein Taghvaei , Yongxin Chen

Sampling from unnormalized densities is analogous to the generative modeling problem, but the target distribution is defined by a known energy function instead of data samples. Because evaluating the energy function is often costly, a…

Machine Learning · Computer Science 2026-05-06 Aaron Havens , Brian Karrer , Neta Shaul

Although generative diffusion models (GDMs) are widely used in practice, their theoretical foundations remain limited, especially concerning the impact of different discretization schemes applied to the underlying stochastic differential…

Numerical Analysis · Mathematics 2026-01-27 Emanuel Pfarr , Radu Timofte , Frank Werner