Related papers: Expressivity of Transformers: A Tropical Geometry …
We introduce a simple, easy to implement, and computationally efficient tropical convolutional neural network architecture that is robust against adversarial attacks. We exploit the tropical nature of piece-wise linear neural networks by…
Tropical geometry gives a bound on the ranks of divisors on curves in terms of the combinatorics of the dual graph of a degeneration. We show that for a family of examples, curves realizing this bound might only exist over certain…
Vision Transformers have made remarkable progress in recent years, achieving state-of-the-art performance in most vision tasks. A key component of this success is due to the introduction of the Multi-Head Self-Attention (MHSA) module, which…
In this paper we bring together tropical linear algebra and convex 3-dimensional bodies. We show how certain convex 3-dimensional bodies having 20 vertices and 12 facets can be encoded in a $4\times 4$ integer zero-diagonal matrix $A$. A…
We define nondegenerate tropical complete intersections imitating the corresponding definition in complex algebraic geometry. As in the complex situation, all nonzero intersection multiplicity numbers between tropical hypersurfaces defining…
Transformers are extremely successful machine learning models whose mathematical properties remain poorly understood. Here, we rigorously characterize the behavior of transformers with hardmax self-attention and normalization sublayers as…
We study some basic algorithmic problems concerning the intersection of tropical hypersurfaces in general dimension: deciding whether this intersection is nonempty, whether it is a tropical variety, and whether it is connected, as well as…
For a complex hypersurface of dimension $d \geq 1$ in a toric variety, we construct lifts of tropical $(p, q)$-cycles with $p+q=d$ in the associated tropical hypersurface. The tropical cycles we consider are described by Minkowski weights,…
We investigate topological properties of density matrices motivated by the question to what extent phenomena like topological insulators and superconductors can be generalized to mixed states in the framework of open quantum systems. The…
We investigate finite-temperature observables in three-dimensional large $N$ critical vector models taking into account the effects suppressed by $1\over N$. Such subleading contributions are captured by the fluctuations of the…
An unconstrained optimization problem is formulated in terms of tropical mathematics to minimize a functional that is defined on a vector set by a matrix and calculated through multiplicative conjugate transposition. For some particular…
Tropical geometry has recently found several applications in the analysis of neural networks with piecewise linear activation functions. This paper presents a new look at the problem of tropical polynomial division and its application to…
Transformers have become a standard neural network architecture for many NLP problems, motivating theoretical analysis of their power in terms of formal languages. Recent work has shown that transformers with hard attention are quite…
The Fr\'{e}chet mean is a fundamental notion of central tendency defined as a minimizer of a sum of squared distances in a general metric space. In this paper, we study Fr\'{e}chet means in tropical geometry -- a piecewise linear,…
As transformers are equivariant to the permutation of input tokens, encoding the positional information of tokens is necessary for many tasks. However, since existing positional encoding schemes have been initially designed for NLP tasks,…
Transformer models encounter challenges in scaling hidden dimensions efficiently, as uniformly increasing them inflates computational and memory costs while failing to emphasize the most relevant features for each token. For further…
We express the beta invariant of a loopless matroid as tropical self-intersection number of the diagonal of its matroid fan (a "local" Poincar\'e-Hopf theorem). This provides another example of uncovering the "geometry" of matroids by…
Self-attention mechanisms have revolutionised deep learning architectures, yet their core mathematical structures remain incompletely understood. In this work, we develop a category-theoretic framework focusing on the linear components of…
Learning reduced descriptions of chaotic many-body dynamics is fundamentally challenging: although microscopic equations are Markovian, collective observables exhibit strong memory and exponential sensitivity to initial conditions and…
Transformer-based language models display impressive reasoning-like behavior, yet remain brittle on tasks that require stable symbolic manipulation. This paper develops a unified perspective on these phenomena by interpreting self-attention…