English
Related papers

Related papers: Expressivity of Transformers: A Tropical Geometry …

200 papers

This paper presents a mathematical interpretation of self-attention by connecting it to distributional semantics principles. We show that self-attention emerges from projecting corpus-level co-occurrence statistics into sequence context.…

Machine Learning · Computer Science 2025-11-19 Nihal Mehta

For tropical $n$-variable polynomials $f, g$ a criterion of containment for tropical hypersurfaces $Trop(f)\subset Trop(g)$ is provided in terms of their Newton polyhedra $N(f), N(g)\subset \mathbb{R}^{n+1}$. Namely, $Trop(f)\subset…

Algebraic Geometry · Mathematics 2024-03-04 Dima Grigoriev

Using tropical geometry one can translate problems in enumerative geometry to combinatorial problems. Thus tropical geometry is a powerful tool in enumerative geometry over the complex and real numbers. Results from $\mathbb{A}^1$-homotopy…

Algebraic Geometry · Mathematics 2024-09-27 Andrés Jaramillo Puentes , Sabrina Pauli

We study the capacity of the self-attention key-query channel: for a fixed budget, how many distinct token-token relations can a single layer reliably encode? We introduce Relational Graph Recognition, where the key-query channel encodes a…

Machine Learning · Computer Science 2026-02-04 Micah Adler

Expressivity plays a fundamental role in evaluating deep neural networks, and it is closely related to understanding the limit of performance improvement. In this paper, we propose a three-pipeline training framework based on critical…

Machine Learning · Computer Science 2020-12-17 Gege Zhang

In tropical geometry, one studies algebraic curves using combinatorial techniques via the tropicalization procedure. The tropicalization depends on a map to an algebraic torus and the combinatorial methods are most useful when the…

Algebraic Geometry · Mathematics 2022-12-07 Trevor Gunn , Philipp Jell

We study the behavior of phylogenetic tree shapes in the tropical geometric interpretation of tree space. Tree shapes are formally referred to as tree topologies; a tree topology can also be thought of as a tree combinatorial type, which is…

Combinatorics · Mathematics 2023-01-25 Bo Lin , Anthea Monod , Ruriko Yoshida

We present TropNNC, a framework for compressing neural networks with linear and convolutional layers and ReLU activations using tropical geometry. By representing a network's output as a tropical rational function, TropNNC enables…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Konstantinos Fotopoulos , Petros Maragos , Panagiotis Misiakos

The Transformer architecture has achieved tremendous success in natural language processing, computer vision, and scientific computing through its self-attention mechanism. However, its core components-positional encoding and attention…

Machine Learning · Computer Science 2025-11-13 Xianshuai Shi , Jianfeng Zhu , Leibo Liu

Compressibility transformations are used to relate hypersonic zero-pressure-gradient (ZPG) turbulent boundary layers (TBLs) to incompressible reference states, but their assessment has largely focused on the collapse of transformed mean…

Fluid Dynamics · Physics 2025-12-11 Engin Danis

Transformers serve as the foundation of most modern large language models. To mitigate the quadratic complexity of standard full attention, various efficient attention mechanisms, such as linear and hybrid attention, have been developed. A…

Machine Learning · Computer Science 2026-02-03 Xiaowei Ye , Xiaoyu He , Chao Liao , Chen Wu , Pinyan Lu

Measure-theoretic slow entropy is a more refined invariant than the classical measure-theoretic entropy to characterize the complexity of dynamical systems with subexponential growth rates of distinguishable orbit types. In this paper we…

Dynamical Systems · Mathematics 2021-09-20 Shilpak Banerjee , Philipp Kunde , Daren Wei

In deep learning theory, the covariance matrix of the representations serves as a proxy to examine the network's trainability. Motivated by the success of Transformers, we study the covariance matrix of a modified Softmax-based attention…

Machine Learning · Statistics 2023-12-12 Lorenzo Noci , Chuning Li , Mufan Bill Li , Bobby He , Thomas Hofmann , Chris Maddison , Daniel M. Roy

We consider the response of a multicomponent body to $n$ fields, such as electric fields, magnetic fields, temperature gradients, concentration gradients, etc., where each component, which is possibly anisotropic, may cross couple the…

Materials Science · Physics 2016-02-23 Mordehai Milgrom , Graeme W. Milton

The transformer has revolutionized modern AI across language, vision, and beyond. It consists of $L$ layers, each running $H$ attention heads in parallel and feeding the combined output to the subsequent layer. In attention, the input…

Computational Complexity · Computer Science 2026-03-13 Barna Saha , Yinzhan Xu , Christopher Ye , Hantao Yu

Under real-analytic assumptions on decoder-only Transformers, recent work shows that the map from discrete prompts to last-token hidden states is generically injective on finite prompt sets. We refine this picture: for each layer $\ell$ we…

Machine Learning · Computer Science 2025-11-20 Mikael von Strauss

We study algebraic and combinatorial aspects of (classical) projections of $m$-dimensional tropical varieties onto $(m+1)$-dimensional planes. Building upon the work of Sturmfels, Tevelev, and Yu on tropical elimination as well as the work…

Algebraic Geometry · Mathematics 2010-04-23 Kerstin Hept , Thorsten Theobald

We study the geometry of tropical extensions of hyperfields, including the ordinary, signed and complex tropical hyperfields. We introduce the framework of 'enriched valuations' as hyperfield homomorphisms to tropical extensions, and show…

Algebraic Geometry · Mathematics 2024-11-27 James Maxwell , Ben Smith

Transformers, which are state-of-the-art in most machine learning tasks, represent the data as sequences of vectors called tokens. This representation is then exploited by the attention function, which learns dependencies between tokens and…

Machine Learning · Computer Science 2025-01-31 Valérie Castin , Pierre Ablin , José Antonio Carrillo , Gabriel Peyré

The transformer is the most popular neural architecture for language modeling. The cornerstone of the transformer is its global attention mechanism, which lets the model aggregate information from all preceding tokens before generating the…

Computation and Language · Computer Science 2026-05-20 Jiaoda Li , Ryan Cotterell
‹ Prev 1 3 4 5 6 7 10 Next ›