Related papers: Geometric and Spectral Alignment for Deep Neural N…
Local models are schemes defined in linear algebra terms that describe the 'etale local structure of integral models for Shimura varieties and other moduli spaces. We point out that the flatness conjecture of Rapoport-Zink on local models…
Branchwidth determines how graphs, and more generally, arbitrary connectivity (basically symmetric and submodular) functions could be decomposed into a tree-like structure by specific cuts. We develop a general framework for designing…
This work establishes rigorous mathematical foundations connecting spectral graph theory, algebraic geometry, and string theory. We construct a canonical mapping whereby any finite graph \(G\) defines a compact Riemann surface \(X_{G}\)…
The rank of a hierarchically hyperbolic space is the maximal number of unbounded factors in a standard product region. For hierarchically hyperbolic groups, this coincides with the maximal dimension of a quasiflat. Examples for which the…
We prove sandwich theorems and a Tauberian theorem in the space of compact metric measure spaces, endowed with the Gromov-Hausdorff-Prokhorov (GHP) topology. These results hold with respect to a close relative of Gromov's Lipschitz order.…
It has been recently observed in much of the literature that neural networks exhibit a bottleneck rank property: for larger depths, the activation and weights of neural networks trained with gradient-based methods tend to be of…
Overparameterized shallow neural networks admit substantial parameter redundancy: distinct parameter vectors may represent the same predictor due to hidden-unit permutations, rescalings, and related symmetries. As a result, geometric…
We discuss the problem of determining the complete weight hierarchy of linear error correcting codes associated to Grassmann varieties and, more generally, to Schubert varieties in Grassmannians. The problem is partially solved in the case…
We compare the marked length spectra of isometric actions of groups with non-positively curved features. Inspired by the recent works of Butt we study approximate versions of marked length spectrum rigidity. We show that for pairs of…
We investigate a hierarchy of semidefinite bounds $\vartheta^{(r)}(G)$ for the stability number $\alpha(G)$ of a graph $G$, based on its copositive programming formulation and introduced by de Klerk and Pasechnik [{\em SIAM J. Optim.} 12…
Let X be a smooth projective curve of genus g>1 defined over an algebraically closed field k of characteristic p>0. Let M_X(r) be the moduli space of semi-stable rank r vector bundles with fixed trivial determinant. The relative Frobenius…
Given a finite simple graph G, let G' be its barycentric refinement: it is the graph in which the vertices are the complete subgraphs of G and in which two such subgraphs are connected, if one is contained into the other. If L(0)=0<L(1) <=…
Several recent trends in machine learning theory and practice, from the design of state-of-the-art Gaussian Process to the convergence analysis of deep neural nets (DNNs) under stochastic gradient descent (SGD), have found it fruitful to…
Sparse models for high-dimensional linear regression and machine learning have received substantial attention over the past two decades. Model selection, or determining which features or covariates are the best explanatory variables, is…
For a smooth map between noetherian schemes, Verdier relates the top relative differentials of the map with the twisted inverse image functor `upper shriek'. We show that the associated traces for smooth proper maps can be rendered concrete…
Probabilistic graphical models offer a powerful framework to account for the dependence structure between variables, which is represented as a graph. However, the dependence between variables may render inference tasks intractable. In this…
Recently proposed Gated Linear Networks present a tractable nonlinear network architecture, and exhibit interesting capabilities such as learning with local error signals and reduced forgetting in sequential learning. In this work, we…
Deep neural networks have been demonstrated to achieve phenomenal success in many domains, and yet their inner mechanisms are not well understood. In this paper, we investigate the curvature of image manifolds, i.e., the manifold deviation…
Normal matrices, or matrices which commute with their adjoints, are of fundamental importance in pure and applied mathematics. In this paper, we study a natural functional on the space of square complex matrices whose global minimizers are…
A well-conditioned Jacobian spectrum has a vital role in preventing exploding or vanishing gradients and speeding up learning of deep neural networks. Free probability theory helps us to understand and handle the Jacobian spectrum. We…