English

Deep Network with Approximation Error Being Reciprocal of Width to Power of Square Root of Depth

Machine Learning 2021-03-30 v6 Numerical Analysis Numerical Analysis Machine Learning

Abstract

A new network with super approximation power is introduced. This network is built with Floor (x\lfloor x\rfloor) or ReLU (max{0,x}\max\{0,x\}) activation function in each neuron and hence we call such networks Floor-ReLU networks. For any hyper-parameters NN+N\in\mathbb{N}^+ and LN+L\in\mathbb{N}^+, it is shown that Floor-ReLU networks with width max{d,5N+13}\max\{d,\, 5N+13\} and depth 64dL+364dL+3 can uniformly approximate a H\"older function ff on [0,1]d[0,1]^d with an approximation error 3λdα/2NαL3\lambda d^{\alpha/2}N^{-\alpha\sqrt{L}}, where α(0,1]\alpha \in(0,1] and λ\lambda are the H\"older order and constant, respectively. More generally for an arbitrary continuous function ff on [0,1]d[0,1]^d with a modulus of continuity ωf()\omega_f(\cdot), the constructive approximation rate is ωf(dNL)+2ωf(d)NL\omega_f(\sqrt{d}\,N^{-\sqrt{L}})+2\omega_f(\sqrt{d}){N^{-\sqrt{L}}}. As a consequence, this new class of networks overcomes the curse of dimensionality in approximation power when the variation of ωf(r)\omega_f(r) as r0r\to 0 is moderate (e.g., ωf(r)rα\omega_f(r) \lesssim r^\alpha for H\"older continuous functions), since the major term to be considered in our approximation rate is essentially d\sqrt{d} times a function of NN and LL independent of dd within the modulus of continuity.

Keywords

Cite

@article{arxiv.2006.12231,
  title  = {Deep Network with Approximation Error Being Reciprocal of Width to Power of Square Root of Depth},
  author = {Zuowei Shen and Haizhao Yang and Shijun Zhang},
  journal= {arXiv preprint arXiv:2006.12231},
  year   = {2021}
}