Related papers: Sparse juntas on the biased hypercube
To better understand complexity in neural networks, we theoretically investigate the idealised phenomenon of lossless network compressibility, whereby an identical function can be implemented with fewer hidden units. In the setting of…
We find uniform with respect to parameter $p \ (1\leq p\leq\infty)$ upper estimations of best approximations by trigonometric polynomials of classes $C^{\psi}_{\beta,p}$ of periodic functions generated by sequences $\psi(k)$, that decrease…
Universal approximation theorems establish the expressive capacity of neural network architectures. For dynamical systems, existing results are limited to finite time horizons or systems with a globally stable equilibrium, leaving…
This paper studies the approximation property of ReLU neural networks (NNs) to piecewise constant functions with unknown interfaces in bounded regions in $\mathbb{R}^d$. Under the assumption that the discontinuity interface $\Gamma$ may be…
The parameters of a neural network are naturally organized in groups, some of which might not contribute to its overall performance. To prune out unimportant groups of parameters, we can include some non-differentiable penalty to the…
Consider a sequence of Riemannian manifolds $(M^n_i,g_i)$ with scalar curvatures and entropies bounded below by small constants $R_i,\mu_i \geq-\epsilon_i$. The goal of this paper is to understand notions of convergence and the structure of…
We show that any Littlewood--Paley square function $S$ satisfying a minimal local testing condition is dominated by a sparse form, \begin{equation*} \langle (Sf)^2,g \rangle\le C \sum_{I \in \mathscr{S}} \langle \lvert f\rvert\rangle_I^2…
We develop a general theoretical and algorithmic framework for sparse approximation and structured prediction in $\mathcal{P}_2(\Omega)$ with Wasserstein barycenters. The barycenters are sparse in the sense that they are computed from an…
This work establishes that sparse Bayesian neural networks achieve optimal posterior contraction rates over anisotropic Besov spaces and their hierarchical compositions. These structures reflect the intrinsic dimensionality of the…
We consider classes of Boolean functions stable under compositions both from the right and from the left with clones. Motivated by the question how many properties of Boolean functions can be defined by means of linear equations, we focus…
Multivariate functions are typically governed by anisotropic features such as edges in images or shock fronts in solutions of transport-dominated equations. One major goal both for the purpose of compression as well as for an efficient…
We analyze the Fourier growth, i.e. the $L_1$ Fourier weight at level $k$ (denoted $L_{1,k}$), of various well-studied classes of "structured" $\mathbb{F}_2$-polynomials. This study is motivated by applications in pseudorandomness, in…
Confidence bands are confidence sets for an unknown function f, containing all functions within some sup-norm distance of an estimator. In the density estimation, regression, and white noise models, we consider the problem of constructing…
The blow-up lemma states that a system of super-regular pairs contains all bounded degree spanning graphs as subgraphs that embed into a corresponding system of complete pairs. This lemma has far-reaching applications in extremal…
The approximate non-deterministic degree of a Boolean function $f$, denoted $\mathsf{ndeg}_\epsilon(f)$ (written $\mathsf{N}_\epsilon(f)$ for brevity), is the minimum degree of a real polynomial $p$ such that $0 \le |p(x)| \le \epsilon$…
Deep neural networks have emerged as powerful tools for learning operators defined over infinite-dimensional function spaces. However, existing theories frequently encounter difficulties related to dimensionality and limited…
There were established the exact-order estimations of the best uniform approximations by{\psi} the trigonometrical polynoms on the $C^{\psi}_{\beta,p}$ classes of $2\pi$-periodic continuous functions $f$, which are defined by the…
This is a conitunation of [1] and [2]. We prove that if function $f$ belongs to the class $\Lambda_{\omega} \overset{\text{def}}{=} \{f: \omega_{f}(\delta)\leq \text{const} \omega(\delta)\} $ for an arbitrary modulus of continuity $\omega$,…
We show in this work that homology in degree d of a congruence group, in a very general framework, defines a weakly polynomial functor of degree at most 2d and we describe this functor modulo polynomial functors of smaller degree. Our main…
We study Fourier-sparse Boolean functions over general finite Abelian groups. A Boolean function $f : G \to \{-1,+1\}$ is $s$-sparse if it has at most $s$ non-zero Fourier coefficients. We introduce a general notion of granularity of…