English

Optimal Denoising in Score-Based Generative Models: The Role of Data Regularity

Machine Learning 2026-03-18 v3 Machine Learning

Abstract

Score-based generative models achieve state-of-the-art sampling performance by denoising a distribution perturbed by Gaussian noise. In this paper, we focus on a single deterministic denoising step, and compare the optimal denoiser for the quadratic loss, we name ''full-denoising'', to the alternative ''half-denoising'' introduced by Hyv{\"a}rinen (2025). We show that looking at the performance in terms of distance between distributions tells a more nuanced story, with different assumptions on the data leading to very different conclusions. We prove that half-denoising is better than full-denoising for regular enough densities, while full-denoising is better for singular densities such as mixtures of Dirac measures or densities supported on a low-dimensional subspace. In the latter case, we prove that full-denoising can alleviate the curse of dimensionality under a linear manifold hypothesis.

Keywords

Cite

@article{arxiv.2503.12966,
  title  = {Optimal Denoising in Score-Based Generative Models: The Role of Data Regularity},
  author = {Eliot Beyler and Francis Bach},
  journal= {arXiv preprint arXiv:2503.12966},
  year   = {2026}
}
R2 v1 2026-06-28T22:23:17.650Z