Related papers: Geometric and Spectral Alignment for Deep Neural N…
Using Green's hyperplane restriction theorem, we prove that the rank of a Hermitian form on the space of holomorphic polynomials is bounded by a constant depending only on the maximum rank of the form restricted to affine manifolds. As an…
This work provides closed-form solutions and minimum achievable errors for a large class of low-rank approximation problems in Hilbert spaces. The proposed theorem generalizes to the case of bounded linear operators the previous results…
The aim of the present paper is to lay the foundation for a theory of Ehresmann structures in positive characteristic, generalizing the Frobenius-projective and Frobenius-affine structures defined in the previous work. This theory deals…
Alignment, the tendency of adjacent weight matrices in deep networks to develop compatible subspace orientations, underlies gradient flow, Neural Collapse, and representation similarity across architectures. Despite extensive empirical…
Grothendieck-Verdier categories (also known as $\ast$-autonomous categories) generalize rigid monoidal categories, with notable representation-theoretic examples including categories of bimodules, modules over Hopf algebroids, and modules…
Rank-constrained matrix problems appear frequently across science and engineering. The convergence analysis of iterative algorithms developed for these problems often hinges on local error bounds, which correlate the distance to the…
Deep neural networks are highly expressive machine learning models with the ability to interpolate arbitrary datasets. Deep nets are typically optimized via first-order methods and the optimization process crucially depends on the…
Let $G$ be a simple algebraic group of type $A$ or $D$ defined over $\C$ and $T$ be a maximal torus of $G$. For a dominant coweight $\lambda$ of $G$, the $T$-fixed point subscheme $(\bar{Gr}_G^\lambda)^T$ of the Schubert variety…
A recent line of work has established intriguing connections between the generalization/compression properties of a deep neural network (DNN) model and the so-called layer weights' stable ranks. Intuitively, the latter are indicators of the…
Deep ReLU neural networks admit nontrivial functional symmetries: vastly different architectures and parameters (weights and biases) can realize the same function. We address the complete identification problem -- given a function f,…
Many inference problems in structured prediction can be modeled as maximizing a score function on a space of labels, where graphs are a natural representation to decompose the total score into a sum of unary (nodes) and pairwise (edges)…
This paper addresses the following question of neural network identifiability: Does the input-output map realized by a feed-forward neural network with respect to a given nonlinearity uniquely specify the network architecture, weights, and…
We study pairs of non-constant maps between two integral schemes of finite type over two (possibly different) fields of positive characteristic. When the target is quasi-affine, Tamagawa showed that the two maps are equal up to a power of…
We establish structure results for Frobenius kernels of automorphism group schemes for surfaces of general type in positive characteristics. It turns out that there are surprisingly few possibilities. This relies on properties of the famous…
Area-preserving maps have been observed to undergo a universal period-doubling cascade, analogous to the famous Feigenbaum-Coullet-Tresser period doubling cascade in one-dimensional dynamics. A renormalization approach has been used by…
We study the geometry of equivariant, proper maps from homogeneous bundles $G\times_P V$ over flag varieties $G/P$ to representations of $G$, called collapsing maps. Kempf showed that, provided the bundle is completely reducible, the image…
We present a unified theoretical framework connecting the first property of Deep Neural Collapse (DNC1) to the emergence of implicit low-rank bias in nonlinear networks trained with $L^2$ weight decay regularization. Our main contributions…
The Bures--Wasserstein geometry of covariance matrices provides a canonical distance on the statistical manifold of centred Gaussian measures and lies at the intersection of information geometry, quantum information, and optimal transport.…
Exact solution of hard combinatorial optimization problems often relies on strong convex relaxations, but solving these relaxations repeatedly inside a branch-and-bound algorithm can be prohibitively expensive. Hence, we consider this…
Let $X$ be a smooth proper genus 2 curve over an algebraically closed field of characteristic 2. The absolute Frobenius induces a rational map $F$ on the the moduli space $M\_X$ of semi-stable rank 2 vector bundles over $X$, which is…