English

Monge-Kantorovich Fitting With Sobolev Budgets

Machine Learning 2025-04-01 v2 Analysis of PDEs

Abstract

Given m<nm < n, we consider the problem of ``best'' approximating an n-dn\text{-d} probability measure ρ\rho via an m-dm\text{-d} measure ν\nu such that supp ν\mathrm{supp}\ \nu has bounded total ``complexity.'' When ρ\rho is concentrated near an m-dm\text{-d} set we may interpret this as a manifold learning problem with noisy data. However, we do not restrict our analysis to this case, as the more general formulation has broader applications. We quantify ν\nu's performance in approximating ρ\rho via the Monge-Kantorovich (also called Wasserstein) pp-cost Wpp(ρ,ν)\mathbb{W}_p^p(\rho, \nu), and constrain the complexity by requiring supp ν\mathrm{supp}\ \nu to be coverable by an f:RmRnf : \mathbb{R}^{m} \to \mathbb{R}^{n} whose Wk,qW^{k,q} Sobolev norm is bounded by 0\ell \geq 0. This allows us to reformulate the problem as minimizing a functional Jp(f)\mathscr J_p(f) under the Sobolev ``budget'' \ell. This problem is closely related to (but distinct from) principal curves with length constraints when m=1,k=1m=1, k = 1 and an unsupervised analogue of smoothing splines when k>1k > 1. New challenges arise from the higher-order differentiability condition. We study the ``gradient'' of Jp\mathscr J_p, which is given by a certain vector field that we call the barycenter field, and use it to prove a nontrivial (almost) strict monotonicity result. We also provide a natural discretization scheme and establish its consistency. We use this scheme as a toy model for a generative learning task, and by analogy, propose novel interpretations for the role regularization plays in improving training.

Keywords

Cite

@article{arxiv.2409.16541,
  title  = {Monge-Kantorovich Fitting With Sobolev Budgets},
  author = {Forest Kobayashi and Jonathan Hayase and Young-Heon Kim},
  journal= {arXiv preprint arXiv:2409.16541},
  year   = {2025}
}

Comments

Expanded abstract and {\S}6; added conclusion ({\S}7); minor correction to implementation of constraint gradient in {\S}5.3.2; removed unused references; misc typo corrections. 69 pages, 51 pages without figures

R2 v1 2026-06-28T18:55:57.736Z