Near-Optimality of Contrastive Divergence Algorithms
Machine Learning
2025-10-16 v1 Machine Learning
Abstract
We perform a non-asymptotic analysis of the contrastive divergence (CD) algorithm, a training method for unnormalized models. While prior work has established that (for exponential family distributions) the CD iterates asymptotically converge at an rate to the true parameter of the data distribution, we show, under some regularity assumptions, that CD can achieve the parametric rate . Our analysis provides results for various data batching schemes, including the fully online and minibatch ones. We additionally show that CD can be near-optimal, in the sense that its asymptotic variance is close to the Cram\'er-Rao lower bound.
Keywords
Cite
@article{arxiv.2510.13438,
title = {Near-Optimality of Contrastive Divergence Algorithms},
author = {Pierre Glaser and Kevin Han Huang and Arthur Gretton},
journal= {arXiv preprint arXiv:2510.13438},
year = {2025}
}
Comments
54 pages