Implicit Bias of the JKO Scheme
Abstract
Wasserstein gradient flow provides a general framework for minimizing an energy functional over the space of probability measures on a Riemannian manifold . Its canonical time-discretization, the Jordan-Kinderlehrer-Otto (JKO) scheme, produces for any step size a sequence of probability distributions that approximate to first order in Wasserstein gradient flow on . But the JKO scheme also has many other remarkable properties not shared by other first order integrators, e.g. it preserves energy dissipation and exhibits unconditional stability for -geodesically convex functionals . To better understand the JKO scheme we characterize its implicit bias at second order in . We show that are approximated to order by Wasserstein gradient flow on a modified energy obtained by subtracting from the squared metric curvature of times . The JKO scheme therefore adds at second order in a deceleration in directions where the metric curvature of is rapidly changing. This corresponds to canonical implicit biases for common functionals: for entropy the implicit bias is the Fisher information, for KL-divergence it is the Fisher-Hyv{\"a}rinen divergence, and for Riemannian gradient descent it is the kinetic energy in the metric . To understand the differences between minimizing and we study JKO-Flow, Wasserstein gradient flow on , in several simple numerical examples. These include exactly solvable Langevin dynamics on the Bures-Wasserstein space and Langevin sampling from a quartic potential in 1D.
Cite
@article{arxiv.2511.14827,
title = {Implicit Bias of the JKO Scheme},
author = {Peter Halmos and Boris Hanin},
journal= {arXiv preprint arXiv:2511.14827},
year = {2026}
}