English

Near-Optimality of Contrastive Divergence Algorithms

Machine Learning 2025-10-16 v1 Machine Learning

Abstract

We perform a non-asymptotic analysis of the contrastive divergence (CD) algorithm, a training method for unnormalized models. While prior work has established that (for exponential family distributions) the CD iterates asymptotically converge at an O(n1/3)O(n^{-1 / 3}) rate to the true parameter of the data distribution, we show, under some regularity assumptions, that CD can achieve the parametric rate O(n1/2)O(n^{-1 / 2}). Our analysis provides results for various data batching schemes, including the fully online and minibatch ones. We additionally show that CD can be near-optimal, in the sense that its asymptotic variance is close to the Cram\'er-Rao lower bound.

Keywords

Cite

@article{arxiv.2510.13438,
  title  = {Near-Optimality of Contrastive Divergence Algorithms},
  author = {Pierre Glaser and Kevin Han Huang and Arthur Gretton},
  journal= {arXiv preprint arXiv:2510.13438},
  year   = {2025}
}

Comments

54 pages

R2 v1 2026-07-01T06:38:44.569Z