Related papers: Erzeugunsgrad, VC-Dimension and Neural Networks wi…
We prove an exponential separation in sample complexity between Euclidean and hyperbolic representations for learning on hierarchical data under standard Lipschitz regularization. For depth-$R$ hierarchies with branching factor $m$, we…
We introduce "representative generation," extending the theoretical framework for generation proposed by Kleinberg et al. (2024) and formalized by Li et al. (2024), to additionally address diversity and bias concerns in generative models.…
Universal approximation theorems provide a mathematical explanation for the expressive power of neural networks. They assert that, under mild conditions on the activation function, feedforward neural networks are dense in broad function…
Reasoning is a fundamental problem for computers and deeply studied in Artificial Intelligence. In this paper, we specifically focus on answering multi-hop logical queries on Knowledge Graphs (KGs). This is a complicated task because, in…
Despite recent advances in representation learning in hypercomplex (HC) space, this subject is still vastly unexplored in the context of graphs. Motivated by the complex and quaternion algebras, which have been found in several contexts to…
Responding to Hodel et al.'s (2024) call for a formal definition of task relatedness in re-arc, we present the first 9-category taxonomy of all 400 tasks, validated at 97.5% accuracy via rule-based code analysis. We prove the taxonomy's…
We consider for $d\geq 1$ the graded commutative $\mathbb{Q}$-algebra $\mathcal{A}(d):=H^*(\operatorname{Hilb}^d(\mathbb{C}^2);\mathbb{Q})$, which is also connected to the study of generalised Hurwitz spaces by work of the first author.…
Lance Bryant noticed in his thesis that there was a flaw in our paper "Associated graded rings of one-dimensional analytically irreducible rings", J. Algebra 304 (2006), 349-358. It can be fixed by adding a condition, called the BF…
Knowledge Graph Embedding models, representing entities and edges in a low-dimensional space, have been extremely successful at solving tasks related to completing and exploring Knowledge Graphs (KGs). One of the key aspects of training…
This work addressed the problem of learning a network with communication between vertices. The communication between vertices is presented in the form of perturbation on the measure. We studied the scenario where samples are drawn from a…
We compute the generating function of column-strict plane partitions with parts in {1,2,...,n}, at most c columns, p rows of odd length and k parts equal to n. This refines both, Krattenthaler's ["The major counting of nonintersecting…
Let $G$ be a simple, simply connected algebraic group over an algebraically closed field of positive characteristic $p$. In recent work, the authors have studied a graded analogue of the category of rational $G$-modules. These gradings are…
We consider a deep matrix factorization model of covariance matrices trained with the Bures-Wasserstein distance. While recent works have made advances in the study of the optimization problem for overparametrized low-rank matrix…
Let $X$ be a complete variety of dimension $n$ over an algebraically closed field $\mathbf{K}$. Let $V_\bullet$ be a graded linear series associated to a line bundle $L$ on $X$, that is, a collection $\{V_m\}_{m\in\mathbb{N}}$ of vector…
Physics-informed neural networks offered an alternate way to solve several differential equations that govern complicated physics. However, their success in predicting the acoustic field is limited by the vanishing-gradient problem that…
We state a conjecture (due to M. Duflo) analogous to the Kashiwara--Vergne conjecture in the case of a characteristic $p>2$, where the role of the Campbell--Hausdorff series is played by the Jacobson element. We prove a simpler version of…
The exploding and vanishing gradient problem has been the major conceptual principle behind most architecture and training improvements in recurrent neural networks (RNNs) during the last decade. In this paper, we argue that this principle,…
Recently, the authors of \cite{SYZ22} developed a neural network with width $36d(2d + 1)$ and depth $11$, which utilizes a special activation function called the elementary universal activation function, to achieve the super approximation…
We show that bounded type implies finite type for a constructible subcategory of the module category of a finitely generated algebra over a field, which is a variant of the first Brauer-Thrall conjecture. A full subcategory is constructible…
Generating realistic graph-structured data is challenging due to discrete structures, variable sizes, and class-specific connectivity patterns that resist conventional generative modelling. While recent graph generation methods employ…