English

A unified perspective on fine-tuning and sampling with diffusion and flow models

Machine Learning 2026-05-04 v1 Machine Learning Optimization and Control

Abstract

We study the problem of training diffusion and flow generative models to sample from target distributions defined by an exponential tilting of a base density; a formulation that subsumes both sampling from unnormalized densities and reward fine-tuning of pre-trained models. This problem can be approached from a stochastic optimal control (SOC) perspective, using adjoint-based or score matching methods, or from a non-equilibrium thermodynamics perspective. We provide a unified framework encompassing these approaches and make three main contributions: (i) bias-variance decompositions revealing that Adjoint Matching/Sampling and Novel Score Matching have finite gradient variance, while Target and Conditional Score Matching do not; (ii) norm bounds on the lean adjoint ODE that theoretically support the effectiveness of adjoint-based methods; and (iii) adaptations of the CMCD and NETS loss functions, along with novel Crooks and Jarzynski identities, to the exponential tilting setting. We validate our analysis with reward fine-tuning experiments on Stable Diffusion 1.5 and 3.

Keywords

Cite

@article{arxiv.2605.00229,
  title  = {A unified perspective on fine-tuning and sampling with diffusion and flow models},
  author = {Carles Domingo-Enrich and Yuanqi Du and Michael S. Albergo},
  journal= {arXiv preprint arXiv:2605.00229},
  year   = {2026}
}
R2 v1 2026-07-01T12:44:31.698Z