English

Neural Empirical Bayes

Machine Learning 2020-04-23 v2 Machine Learning

Abstract

We unify kernel density estimation\textit{kernel density estimation} and empirical Bayes\textit{empirical Bayes} and address a set of problems in unsupervised learning with a geometric interpretation of those methods, rooted in the concentration of measure\textit{concentration of measure} phenomenon. Kernel density is viewed symbolically as XYX\rightharpoonup Y where the random variable XX is smoothed to Y=X+N(0,σ2Id)Y= X+N(0,\sigma^2 I_d), and empirical Bayes is the machinery to denoise in a least-squares sense, which we express as XYX \leftharpoondown Y. A learning objective is derived by combining these two, symbolically captured by XYX \rightleftharpoons Y. Crucially, instead of using the original nonparametric estimators, we parametrize the energy function\textit{the energy function} with a neural network denoted by ϕ\phi; at optimality, ϕlogf\nabla \phi \approx -\nabla \log f where ff is the density of YY. The optimization problem is abstracted as interactions of high-dimensional spheres which emerge due to the concentration of isotropic gaussians. We introduce two algorithmic frameworks based on this machinery: (i) a "walk-jump" sampling scheme that combines Langevin MCMC (walks) and empirical Bayes (jumps), and (ii) a probabilistic framework for associative memory\textit{associative memory}, called NEBULA, defined \`{a} la Hopfield by the gradient flow\textit{gradient flow} of the learned energy to a set of attractors. We finish the paper by reporting the emergence of very rich "creative memories" as attractors of NEBULA for highly-overlapping spheres.

Keywords

Cite

@article{arxiv.1903.02334,
  title  = {Neural Empirical Bayes},
  author = {Saeed Saremi and Aapo Hyvarinen},
  journal= {arXiv preprint arXiv:1903.02334},
  year   = {2020}
}

Comments

23 pages, 10 figures

R2 v1 2026-06-23T07:59:46.211Z