English

Kernel Approximation of Fisher-Rao Gradient Flows

Machine Learning 2024-10-29 v1 Machine Learning Analysis of PDEs

Abstract

The purpose of this paper is to answer a few open questions in the interface of kernel methods and PDE gradient flows. Motivated by recent advances in machine learning, particularly in generative modeling and sampling, we present a rigorous investigation of Fisher-Rao and Wasserstein type gradient flows concerning their gradient structures, flow equations, and their kernel approximations. Specifically, we focus on the Fisher-Rao (also known as Hellinger) geometry and its various kernel-based approximations, developing a principled theoretical framework using tools from PDE gradient flows and optimal transport theory. We also provide a complete characterization of gradient flows in the maximum-mean discrepancy (MMD) space, with connections to existing learning and inference algorithms. Our analysis reveals precise theoretical insights linking Fisher-Rao flows, Stein flows, kernel discrepancies, and nonparametric regression. We then rigorously prove evolutionary Γ\Gamma-convergence for kernel-approximated Fisher-Rao flows, providing theoretical guarantees beyond pointwise convergence. Finally, we analyze energy dissipation using the Helmholtz-Rayleigh principle, establishing important connections between classical theory in mechanics and modern machine learning practice. Our results provide a unified theoretical foundation for understanding and analyzing approximations of gradient flows in machine learning applications through a rigorous gradient flow and variational method perspective.

Keywords

Cite

@article{arxiv.2410.20622,
  title  = {Kernel Approximation of Fisher-Rao Gradient Flows},
  author = {Jia-Jie Zhu and Alexander Mielke},
  journal= {arXiv preprint arXiv:2410.20622},
  year   = {2024}
}
R2 v1 2026-06-28T19:37:25.610Z