Related papers: Geometric and Spectral Alignment for Deep Neural N…
Modern statistical learning theory and deep learning characterize generalization primarily in terms of continuous capacity control (e.g., norm-based regularization, margin maximization, low-rank bias). While highly successful in continuous…
This paper studies the quotient geometry of bounded or fixed-rank correlation matrices. We establish a bijection between the set of bounded-rank correlation matrices and a quotient set of a spherical product manifold by an orthogonal group.…
We study regularized deep neural networks (DNNs) and introduce a convex analytic framework to characterize the structure of the hidden layers. We show that a set of optimal hidden layer weights for a norm regularized DNN training problem…
Enforcing orthogonality in neural networks is an antidote for gradient vanishing/exploding problems, sensitivity by adversarial perturbation, and bounding generalization errors. However, many previous approaches are heuristic, and the…
We study the bar-and-joint frameworks in $\mathbb{R}^2$ such that some vertices are constrained to lie on some lines. The generic rigidity of such frameworks is characterised by Streinu and Theran (2010). Katoh and Tanigawa (2013) remarked…
Let $\mathtt{k}$ be an algebraic closure of a finite field $\mathbb{F}_{q}$ of characteristic $p$. Let $G$ be a connected unipotent group over $\mathtt{k}$ equipped with an $\mathbb{F}_q$-structure given by a Frobenius map $F:G\to G$. We…
Universal approximation theorems show that neural networks can approximate any continuous function; however, the number of parameters may grow exponentially with the ambient dimension, so these results do not fully explain the practical…
In this paper we investigate the family of functions representable by deep neural networks (DNN) with rectified linear units (ReLU). We give an algorithm to train a ReLU DNN with one hidden layer to *global optimality* with runtime…
The flat rank of a totally disconnected locally compact group G, denoted flat-rk(G), is an invariant of the topological group structure of G. It is defined thanks to a natural distance on the space of compact open subgroups of G. For a…
We propose a scalable framework for the learning of high-dimensional parametric maps via adaptively constructed residual network (ResNet) maps between reduced bases of the inputs and outputs. When just few training data are available, it is…
Graph neural networks for node classification are typically trained by gradient descent over hundreds or thousands of epochs. Recent work has shown that, when properly tuned, classic GCN/SAGE/GAT architectures can match graph transformers…
Nonlinear kernels can be approximated using finite-dimensional feature maps for efficient risk minimization. Due to the inherent trade-off between the dimension of the (mapped) feature space and the approximation accuracy, the key problem…
In this paper, we develop the foundations of the theory of quasiregular mappings in general metric measure spaces. In particular, nine definitions of quasiregularity for a discrete open mapping with locally bounded multiplicity are proved…
Besides classical feed-forward neural networks such as multilayer perceptrons, also neural ordinary differential equations (neural ODEs) have gained particular interest in recent years. Neural ODEs can be interpreted as an infinite depth…
Scaling factors in residual branches have emerged as a prevalent method for boosting neural network performance, especially in normalization-free architectures. While prior work has primarily examined scaling effects from an optimization…
We analyze the deformation theory of equivariant vector bundles. In particular, we provide an effective criterion for verifying whether all infinitesimal deformations preserve the equivariant structure. As an application, using rigidity of…
Let $G$ be a connected Lie group. In this paper, we study the density of the images of individual power maps $P_k:G\to G:g\mapsto g^k$. We give criteria for the density of $P_k(G)$ in terms of regular elements, as well as Cartan subgroups.…
Previous work of the second author and Wolf showed that given a set $A\subseteq \mathbb{F}_p^n$ of bounded $\textrm{VC}_2$-dimension, there is a high rank quadratic factor $\mathcal{B}$ of bounded complexity such that $A$ is approximately…
We introduce a principled approach for unsupervised structure learning of deep neural networks. We propose a new interpretation for depth and inter-layer connectivity where conditional independencies in the input distribution are encoded…
Let $A$ be an Artinian local ring with algebraically closed residue field $k$, and let $\mathbf{G}$ be an affine smooth group scheme over $A$. The Greenberg functor $\mathcal{F}$ associates to $\mathbf{G}$ a linear algebraic group…