Related papers: A Quantitative Functional Central Limit Theorem fo…
We consider the approximation rates of shallow neural networks with respect to the variation norm. Upper bounds on these rates have been established for sigmoidal and ReLU activation functions, but it has remained an important open problem…
Let $\{X_k\}_{k \in \mathbb{Z}}$ be a stationary Gaussian process with values in a separable Hilbert space $\mathcal{H}_1$, and let $G:\mathcal{H}_1 \to \mathcal{H}_2$ be an operator acting on $X_k$. Under suitable conditions on the…
We illustrate an approach that can be exploited for constructing neural networks which a priori obey physical laws. We start with a simple single-layer neural network (NN) but refrain from choosing the activation functions yet. Under…
In this effort, we derive a formula for the integral representation of a shallow neural network with the ReLU activation function. We assume that the outer weighs admit a finite $L_1$-norm with respect to Lebesgue measure on the sphere. For…
The general model of coagulation is considered. For basic classes of unbounded coagulation kernels the central limit theorem (CLT) is obtained for the fluctuations around the dynamic law of large numbers (LLN). A rather precise rate of…
In this paper we survey the almost sure central limit theorem and its functional form (quenched) for stationary and ergodic processes. For additive functionals of a stationary and ergodic Markov chain these theorems are known under the…
The celebrated universal approximation theorems for neural networks roughly state that any reasonable function can be arbitrarily well-approximated by a network whose parameters are appropriately chosen real numbers. This paper examines the…
We analyze a simple one-hidden-layer neural network with ReLU activation functions and fixed biases, with one-dimensional input and output. We study both continuous and discrete versions of the model, and we rigorously prove the convergence…
We focus on a specific class of shallow neural networks with a single hidden layer, namely those with $L_2$-normalised data and either a sigmoid-shaped Gaussian error function ("erf") activation or a Gaussian Error Linear Unit (GELU)…
Let $W_{\infty}(\beta)$ be the limit of the Biggins martingale $W_n(\beta)$ associated to a supercritical branching random walk with mean number of offspring $m$. We prove a functional central limit theorem stating that as $n\to\infty$ the…
In this paper, we consider regression problems with one-hidden-layer neural networks (1NNs). We distill some properties of activation functions that lead to $\mathit{local~strong~convexity}$ in the neighborhood of the ground-truth…
We estimate linear functionals in the classical deconvolution problem by kernel estimators. We obtain a uniform central limit theorem with $\sqrt{n}$-rate on the assumption that the smoothness of the functionals is larger than the…
Let $(X_i)$ be a stationary and ergodic Markov chain with kernel $Q$, $f$ an $L^2$ function on its state space. If $Q$ is a normal operator and $f = (I-Q)^{1/2}g$ (which is equivalent to the convergence of $\sum_{n=1}^\infty…
Neural networks are widely used to approximate unknown functions in control. A common neural network architecture uses a single hidden layer (i.e. a shallow network), in which the input parameters are fixed in advance and only the output…
Quantitative multivariate central limit theorems for general functionals of possibly non-symmetric and non-homogeneous infinite Rademacher sequences are proved by combining discrete Malliavin calculus with the smart path method for normal…
This paper studies the approximation capacity of neural networks with an arbitrary activation function and with norm constraint on the weights. Upper and lower bounds on the approximation error of these networks are computed for smooth…
In this paper we study shallow neural network functions which are linear combinations of compositions of activation and quadratic functions, replacing standard affine linear functions, often called neurons. We show the universality of this…
Let $Q$ be a transition probability on a measurable space $E$ which admits an invariant probability measure, let $(X_n)_n$ be a Markov chain associated to $Q$, and let $\xi$ be a real-valued measurable function on $E$, and $S_n=\sum…
We study the sample complexity of learning one-hidden-layer convolutional neural networks (CNNs) with non-overlapping filters. We propose a novel algorithm called approximate gradient descent for training CNNs, and show that, with high…
This paper establishes central limit theorems for Polyak-Ruppert averaged Q-learning under asynchronous updates. We prove a non-asymptotic central limit theorem, where the convergence rate in Wasserstein distance explicitly reflects the…