English
Related papers

Related papers: Expressivity of Transformers: A Tropical Geometry …

200 papers

The self-attention mechanism, a cornerstone of Transformer-based state-of-the-art deep learning architectures, is largely heuristic-driven and fundamentally challenging to interpret. Establishing a robust theoretical foundation to explain…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Laziz U. Abdullaev , Maksim Tkachenko , Tan M. Nguyen

Until now, it has been difficult for volumetric super-resolution to utilize the recent advances in transformer-based models seen in 2D super-resolution. The memory required for self-attention in 3D volumes limits the receptive field.…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 August Leander Høeg , Sophia W. Bardenfleth , Hans Martin Kjer , Tim B. Dyrby , Vedrana Andersen Dahl , Anders Dahl

Vision Transformers and their variants have achieved remarkable success in diverse visual perception tasks. Despite their effectiveness, they suffer from two significant limitations. First, the quadratic computational complexity of…

Computer Vision and Pattern Recognition · Computer Science 2026-02-16 Ali K. Rahimian , Manish K. Govind , Subhajit Maity , Dominick Reilly , Christian Kümmerle , Srijan Das , Aritra Dutta

Transformers empirically perform precise probabilistic reasoning in carefully constructed ``Bayesian wind tunnels'' and in large-scale language models, yet the mechanisms by which gradient-based learning creates the required internal…

Machine Learning · Statistics 2026-05-19 Naman Agarwal , Siddhartha R. Dalal , Vishal Misra

We study the geometry of metrics and convexity structures on the space of phylogenetic trees, which is here realized as the tropical linear space of all \ ultrametrics. The ${\rm CAT}(0)$-metric of Billera-Holmes-Vogtman arises from the…

Metric Geometry · Mathematics 2018-02-19 Bo Lin , Bernd Sturmfels , Xiaoxian Tang , Ruriko Yoshida

Given an algebraic variety defined over a discrete valuation field and a skeleton of its Berkovich analytification, the tropicalization process transforms function field of the variety to a semifield of tropical functions on the skeleton.…

Algebraic Geometry · Mathematics 2025-03-27 Omid Amini , Shu Kawaguchi , JuAe Song

In this paper we give an interpretation to the boundary points of the compactification of the parameter space of convex projective structures on an n-manifold M. These spaces are closed semi-algebraic subsets of the variety of characters of…

Geometric Topology · Mathematics 2014-10-01 Daniele Alessandrini

For a class of piecewise hyperbolic maps in two dimensions, we propose a combinatorial definition of topological entropy by counting the maximal, open, connected components of the phase space on which iterates of the map are smooth. We…

Dynamical Systems · Mathematics 2020-03-11 Mark F. Demers

Transformers are built upon multi-head scaled dot-product attention and positional encoding, which aim to learn the feature representations and token dependencies. In this work, we focus on enhancing the distinctive representation by…

Computer Vision and Pattern Recognition · Computer Science 2022-07-12 Litao Yu , Jian Zhang

Transformer self-attention can be interpreted as a gradient flow on the unit sphere, in which tokens evolve under softmax interaction potentials and tend to form clusters. While prior work has established clustering behavior for single-head…

Machine Learning · Computer Science 2026-05-11 Ayan Pendharkar

We re-explore the symmetries of a weakly isolated horizon (WIH) from the perspective of freedom in the choice of intrinsic data. The supertranslations are realized as additional symmetries. Further, it is shown that all smooth vector fields…

General Relativity and Quantum Cosmology · Physics 2021-01-04 Amit Ghosh , Avirup Ghosh , Pritam Nanda

Transformers trained in low precision can suffer forward-error amplification. We give a first-order, module-wise theory that predicts when and where errors grow. For self-attention we derive a per-layer bound that factorizes into three…

Machine Learning · Computer Science 2025-10-28 Jinwoo Baek

The Transformer architecture has revolutionized the field of sequence modeling and underpins the recent breakthroughs in large language models (LLMs). However, a comprehensive mathematical theory that explains its structure and operations…

Machine Learning · Computer Science 2026-04-14 Xue-Cheng Tai , Hao Liu , Lingfeng Li , Raymond H. Chan

We study tropical degree bounds, stable tropical intersections, and tropical B\'ezout-type estimates through the geometry of Newton polytopes, mixed subdivisions, and lattice indices. We establish an upper bound for the tropical degree of a…

Algebraic Geometry · Mathematics 2026-05-26 Mounir Nisse

Window-based transformers have demonstrated outstanding performance in super-resolution tasks due to their adaptive modeling capabilities through local self-attention (SA). However, they exhibit higher computational complexity and inference…

Computer Vision and Pattern Recognition · Computer Science 2024-09-27 Zhenyu Hu , Wanjie Sun

The tangential layers are characterized by a bulk plasma velocity and a magnetic field that are perpendicular to the gradient direction. They have been extensively described in the frame of the Magneto-Hydro-Dynamic (MHD) theory. But the…

Space Physics · Physics 2009-11-10 Fabrice Mottez

Deep learning models have been widely applied in various aspects of daily life. Many variant models based on deep learning structures have achieved even better performances. Attention-based architectures have become almost ubiquitous in…

Machine Learning · Computer Science 2022-02-25 Zhiying Fang , Yidong Ouyang , Ding-Xuan Zhou , Guang Cheng

We implement new techniques involving Artin fans to study the realizability of tropical stable maps in superabundant combinatorial types. Our approach is to understand the skeleton of a fundamental object in logarithmic Gromov--Witten…

Algebraic Geometry · Mathematics 2017-06-27 Dhruv Ranganathan

An inherent challenge in computing fully-explicit generalization bounds for transformers involves obtaining covering number estimates for the given transformer class $T$. Crude estimates rely on a uniform upper bound on the local-Lipschitz…

Machine Learning · Computer Science 2025-02-07 Yannick Limmer , Anastasis Kratsios , Xuwei Yang , Raeid Saqur , Blanka Horvath

Neural networks equipped with self-attention have parallelizable computation, light-weight structure, and the ability to capture both long-range and local dependencies. Further, their expressive power and performance can be boosted by using…

Computation and Language · Computer Science 2019-03-27 Tao Shen , Tianyi Zhou , Guodong Long , Jing Jiang , Chengqi Zhang