Related papers: Expressivity of Transformers: A Tropical Geometry …
The self-attention mechanism, a cornerstone of Transformer-based state-of-the-art deep learning architectures, is largely heuristic-driven and fundamentally challenging to interpret. Establishing a robust theoretical foundation to explain…
Until now, it has been difficult for volumetric super-resolution to utilize the recent advances in transformer-based models seen in 2D super-resolution. The memory required for self-attention in 3D volumes limits the receptive field.…
Vision Transformers and their variants have achieved remarkable success in diverse visual perception tasks. Despite their effectiveness, they suffer from two significant limitations. First, the quadratic computational complexity of…
Transformers empirically perform precise probabilistic reasoning in carefully constructed ``Bayesian wind tunnels'' and in large-scale language models, yet the mechanisms by which gradient-based learning creates the required internal…
We study the geometry of metrics and convexity structures on the space of phylogenetic trees, which is here realized as the tropical linear space of all \ ultrametrics. The ${\rm CAT}(0)$-metric of Billera-Holmes-Vogtman arises from the…
Given an algebraic variety defined over a discrete valuation field and a skeleton of its Berkovich analytification, the tropicalization process transforms function field of the variety to a semifield of tropical functions on the skeleton.…
In this paper we give an interpretation to the boundary points of the compactification of the parameter space of convex projective structures on an n-manifold M. These spaces are closed semi-algebraic subsets of the variety of characters of…
For a class of piecewise hyperbolic maps in two dimensions, we propose a combinatorial definition of topological entropy by counting the maximal, open, connected components of the phase space on which iterates of the map are smooth. We…
Transformers are built upon multi-head scaled dot-product attention and positional encoding, which aim to learn the feature representations and token dependencies. In this work, we focus on enhancing the distinctive representation by…
Transformer self-attention can be interpreted as a gradient flow on the unit sphere, in which tokens evolve under softmax interaction potentials and tend to form clusters. While prior work has established clustering behavior for single-head…
We re-explore the symmetries of a weakly isolated horizon (WIH) from the perspective of freedom in the choice of intrinsic data. The supertranslations are realized as additional symmetries. Further, it is shown that all smooth vector fields…
Transformers trained in low precision can suffer forward-error amplification. We give a first-order, module-wise theory that predicts when and where errors grow. For self-attention we derive a per-layer bound that factorizes into three…
The Transformer architecture has revolutionized the field of sequence modeling and underpins the recent breakthroughs in large language models (LLMs). However, a comprehensive mathematical theory that explains its structure and operations…
We study tropical degree bounds, stable tropical intersections, and tropical B\'ezout-type estimates through the geometry of Newton polytopes, mixed subdivisions, and lattice indices. We establish an upper bound for the tropical degree of a…
Window-based transformers have demonstrated outstanding performance in super-resolution tasks due to their adaptive modeling capabilities through local self-attention (SA). However, they exhibit higher computational complexity and inference…
The tangential layers are characterized by a bulk plasma velocity and a magnetic field that are perpendicular to the gradient direction. They have been extensively described in the frame of the Magneto-Hydro-Dynamic (MHD) theory. But the…
Deep learning models have been widely applied in various aspects of daily life. Many variant models based on deep learning structures have achieved even better performances. Attention-based architectures have become almost ubiquitous in…
We implement new techniques involving Artin fans to study the realizability of tropical stable maps in superabundant combinatorial types. Our approach is to understand the skeleton of a fundamental object in logarithmic Gromov--Witten…
An inherent challenge in computing fully-explicit generalization bounds for transformers involves obtaining covering number estimates for the given transformer class $T$. Crude estimates rely on a uniform upper bound on the local-Lipschitz…
Neural networks equipped with self-attention have parallelizable computation, light-weight structure, and the ability to capture both long-range and local dependencies. Further, their expressive power and performance can be boosted by using…