English

Is RL fine-tuning harder than regression? A PDE learning approach for diffusion models

Machine Learning 2025-09-03 v1 Optimization and Control Probability Statistics Theory Machine Learning Statistics Theory

Abstract

We study the problem of learning the optimal control policy for fine-tuning a given diffusion process, using general value function approximation. We develop a new class of algorithms by solving a variational inequality problem based on the Hamilton-Jacobi-Bellman (HJB) equations. We prove sharp statistical rates for the learned value function and control policy, depending on the complexity and approximation errors of the function class. In contrast to generic reinforcement learning problems, our approach shows that fine-tuning can be achieved via supervised regression, with faster statistical rate guarantees.

Keywords

Cite

@article{arxiv.2509.02528,
  title  = {Is RL fine-tuning harder than regression? A PDE learning approach for diffusion models},
  author = {Wenlong Mou},
  journal= {arXiv preprint arXiv:2509.02528},
  year   = {2025}
}
R2 v1 2026-07-01T05:17:44.215Z