Related papers: Schauder Bases for $C[0, 1]$ Using ReLU, Softplus …
While it is well-known that neural networks enjoy excellent approximation capabilities, it remains a big challenge to compute such approximations from point samples. Based on tools from Information-based complexity, recent work by Grohs and…
Meshfree methods based on radial basis function (RBF) approximation are of interest for numerical solution of partial differential equations (PDEs) because they are flexible with respect to the geometry of the computational domain, they can…
In recent years, functional neural networks have been proposed and studied in order to approximate nonlinear continuous functionals defined on $L^p([-1, 1]^s)$ for integers $s\ge1$ and $1\le p<\infty$. However, their theoretical properties…
We generalize the famous weight basis constructions of the finite-dimensional irreducible representations of $\mathfrak{sl}(n,\mathbb{C})$ obtained by Gelfand and Tsetlin in 1950. Using combinatorial methods, we construct one such basis for…
Recent findings suggest that consecutive layers of neural networks with the ReLU activation function \emph{fold} the input space during the learning process. While many works hint at this phenomenon, an approach to quantify the folding was…
In this article, we prove that the monomials form a basis for the space of holomorphic functions $(\mathcal{H}(Z), \tau_0)$, where $Z$ denotes either the space $c_0\left(\bigoplus^\infty_{i=1}\ell^i_p \right)$ for some $p\in [1, \infty)$,…
We give the first dimension-efficient algorithms for learning Rectified Linear Units (ReLUs), which are functions of the form $\mathbf{x} \mapsto \max(0, \mathbf{w} \cdot \mathbf{x})$ with $\mathbf{w} \in \mathbb{S}^{n-1}$. Our algorithm…
We study the following two related problems. The first is to determine to what error an arbitrary zonoid in $\mathbb{R}^{d+1}$ can be approximated in the Hausdorff distance by a sum of $n$ line segments. The second is to determine optimal…
This paper investigates the approximation properties of shallow neural networks with activation functions that are powers of exponential functions. It focuses on the dependence of the approximation rate on the dimension and the smoothness…
ReLU is widely seen as the default choice for activation functions in neural networks. However, there are cases where more complicated functions are required. In particular, recurrent neural networks (such as LSTMs) make extensive use of…
One means of fitting functions to high-dimensional data is by providing smoothness constraints. Recently, the following smooth function approximation problem was proposed: given a finite set $E \subset \mathbb{R}^d$ and a function $f: E…
A set of $k$ orthonormal bases of $\mathbb C^d$ is called mutually unbiased if $|\langle e,f\rangle |^2 = 1/d$ whenever $e$ and $f$ are basis vectors in distinct bases. A natural question is for which pairs $(d,k)$ there exist~$k$ mutually…
We introduce Lipschitz continuous and $C^{1,1}$ geometric approximation and interpolation methods for sampled bounded uniformly continuous functions over compact sets and over complements of bounded open sets in $\mathbb{R}^n$ by using…
This paper derives a complete set of quadratic constraints (QCs) for the repeated ReLU. The complete set of QCs is described by a collection of matrix copositivity conditions. We also show that only two functions satisfy all QCs in our…
We show that both $(\mathcal{H}(c_{0}),\tau_{\omega})$ and $(\mathcal{H}_{b}(c_{0}),\tau_{b})$ have a monomial Schauder basis.
There has been a growing interest in expressivity of deep neural networks. However, most of the existing work about this topic focuses only on the specific activation function such as ReLU or sigmoid. In this paper, we investigate the…
In this paper, a notion of Schauder equivalence relation $\mathbb R^\mathbb N/L$ is introduced, where $L$ is a linear subspace of $\mathbb R^\mathbb N$ and the unit vectors form a Schauder basis of $L$. The main theorem is to show that the…
We provide a systematic recipe for translating ReLU approximation results to softmax attention mechanism. This recipe covers many common approximation targets. Importantly, it yields target-specific, economic resource bounds beyond…
It has been shown that deep neural networks of a large enough width are universal approximators but they are not if the width is too small. There were several attempts to characterize the minimum width $w_{\min}$ enabling the universal…
We show that $C^0$-fine approximation of convex functions by smooth (or real analytic) convex functions on $\R^d$ is possible in general if and only if $d=1$. Nevertheless, for $d\geq 2$ we give a characterization of the class of convex…