Related papers: Some density theorems in neural network with varia…
Given a probability density $P({\bf x}|{\boldsymbol \lambda})$, where $\bf x$ represents continuous degrees of freedom and $\lambda$ a set of parameters, it is possible to construct a general identity relating expectations of observable…
Invertible neural networks (INNs) are neural network architectures with invertibility by design. Thanks to their invertibility and the tractability of Jacobian, INNs have various machine learning applications such as probabilistic modeling,…
In [Y.~K.~Hu, K.~A.~Kopotun, X.~M.~Yu, Constr. Approx. 2000], the authors have obtained a characterization of best $n$-term piecewise polynomial approximation spaces as real interpolation spaces between $L^p$ and some spaces of bounded…
We establish global universal approximation theorems on spaces of piecewise linear paths, stating that linear functionals of the corresponding signatures are dense with respect to $L^p$- and weighted norms, under an integrability condition…
Variable Muckenhoupt weights are considered in variable exponent Lebesgue spaces. Applications are given for polynomial approximation in these spaces. Boundedness of averaging operator is proved to gain a transference result. Almost all…
We give the first mathematically rigorous justification of the Local Density Approximation in Density Functional Theory. We provide a quantitative estimate on the difference between the grand-canonical Levy-Lieb energy of a given density…
The classical universal approximation (UA) theorem for neural networks establishes mild conditions under which a feedforward neural network can approximate a continuous function $f$ with arbitrary accuracy. A recent result shows that neural…
It is well known that the output of a Neural Network trained to disentangle between two classes has a probabilistic interpretation in terms of the a-posteriori Bayesian probability, provided that a unary representation is taken for the…
Using a transference result, several inequalities of approximation by entire functions of exponential type in $\mathcal{C}(\mathbf{R})$, the class of bounded uniformly continuous functions defined on $\mathbf{R}:=\left( -\infty ,+\infty…
We study the expressivity of deep neural networks. Measuring a network's complexity by its number of connections or by its number of neurons, we consider the class of functions for which the error of best approximation with networks of a…
Recent years have witnessed strong empirical performance of over-parameterized neural networks on various tasks and many advances in the theory, e.g. the universal approximation and provable convergence to global minimum. In this paper, we…
We consider the sum of power weighted nearest neighbor distances in a sample of size n from a multivariate density f of possibly unbounded support. We give various criteria guaranteeing that this sum satisfies a law of large numbers for…
We introduce a class of fully-connected neural networks whose activation functions, rather than being pointwise, rescale feature vectors by a function depending only on their norm. We call such networks radial neural networks, extending…
We propose a testable universality hypothesis, asserting that seemingly disparate neural network solutions observed in the simple task of modular addition are unified under a common abstract algorithm. While prior work interpreted…
Here we introduce a generalization of the exponential sampling series of optical physics and establish pointwise and uniform convergence theorem, also in a quantitative form. Moreover we compare the error of approximation for Mellin…
In variational inference, the benefits of Bayesian models rely on accurately capturing the true posterior distribution. We propose using neural samplers that specify implicit distributions, which are well-suited for approximating complex…
This paper develops fundamental limits of deep neural network learning by characterizing what is possible if no constraints are imposed on the learning algorithm and on the amount of training data. Concretely, we consider Kolmogorov-optimal…
A candidate explanation of the good empirical performance of deep neural networks is the implicit regularization effect of first order optimization methods. Inspired by this, we prove a convergence theorem for nonconvex composite…
Recent advances in statistical inference have significantly expanded the toolbox of probabilistic modeling. Historically, probabilistic modeling has been constrained to (i) very restricted model classes where exact or approximate…
The primary objective of learning methods is generalization. Classic uniform generalization bounds, which rely on VC-dimension or Rademacher complexity, fail to explain the significant attribute that over-parameterized models in deep…