English

Universal Approximation Depth and Errors of Narrow Belief Networks with Discrete Units

Machine Learning 2014-01-30 v2 Machine Learning Probability

Abstract

We generalize recent theoretical work on the minimal number of layers of narrow deep belief networks that can approximate any probability distribution on the states of their visible units arbitrarily well. We relax the setting of binary units (Sutskever and Hinton, 2008; Le Roux and Bengio, 2008, 2010; Mont\'ufar and Ay, 2011) to units with arbitrary finite state spaces, and the vanishing approximation error to an arbitrary approximation error tolerance. For example, we show that a qq-ary deep belief network with L2+qmδ1q1L\geq 2+\frac{q^{\lceil m-\delta \rceil}-1}{q-1} layers of width nm+logq(m)+1n \leq m + \log_q(m) + 1 for some mNm\in \mathbb{N} can approximate any probability distribution on {0,1,,q1}n\{0,1,\ldots,q-1\}^n without exceeding a Kullback-Leibler divergence of δ\delta. Our analysis covers discrete restricted Boltzmann machines and na\"ive Bayes models as special cases.

Keywords

Cite

@article{arxiv.1303.7461,
  title  = {Universal Approximation Depth and Errors of Narrow Belief Networks with Discrete Units},
  author = {Guido F. Montúfar},
  journal= {arXiv preprint arXiv:1303.7461},
  year   = {2014}
}

Comments

19 pages, 5 figures, 1 table