English
Related papers

Related papers: Attention's forward pass and Frank-Wolfe

200 papers

Large Language Models (LLMs) exhibit exceptional proficiency in handling extensive context windows in natural language. Nevertheless, the quadratic scaling of attention computation relative to sequence length creates substantial efficiency…

Machine Learning · Computer Science 2026-01-26 Xiaoyu Li , Yingyu Liang , Zhenmei Shi , Zhao Song , Song Yue , Jiahao Zhang

Transformers are state-of-the-art in a wide range of NLP tasks and have also been applied to many real-world products. Understanding the reliability and certainty of transformer model predictions is crucial for building trustable machine…

Computation and Language · Computer Science 2021-12-28 Jiahuan Pei , Cheng Wang , György Szarvas

We develop a new numerical method for thin plates falling in inviscid fluid that allows for leading-edge vortex shedding. The inclusion of leading-edge shedding restores physical dynamics to vortex-sheet models of falling bodies, and for…

Fluid Dynamics · Physics 2025-10-02 Yu Jun Loo , Silas Alben

Transformer-based models have achieved remarkable success across a wide range of domains, yet our understanding of their training dynamics remains limited. In this work, we identify a recurrent focus-dilution cycle in attention learning and…

Machine Learning · Computer Science 2026-05-05 Zheng-An Chen , Pengxiao Lin , Zhi-Qin John Xu , Tao Luo

Tensor Attention extends traditional attention mechanisms by capturing high-order correlations across multiple modalities, addressing the limitations of classical matrix-based attention. Meanwhile, Rotary Position Embedding…

Machine Learning · Computer Science 2024-12-25 Xiaoyu Li , Yingyu Liang , Zhenmei Shi , Zhao Song , Mingda Wan

We develop an algorithmic framework for solving convex optimization problems using no-regret game dynamics. By converting the problem of minimizing a convex function into an auxiliary problem of solving a min-max game in a sequential…

Machine Learning · Computer Science 2023-02-21 Jun-Kun Wang , Jacob Abernethy , Kfir Y. Levy

We prove that the block-coordinate Frank-Wolfe (BCFW) algorithm converges with state-of-the-art rates in both convex and nonconvex settings under a very mild "block-iterative" assumption. This appears to be the first result on BCFW…

Optimization and Control · Mathematics 2025-12-17 Gábor Braun , Jannis Halbey , Sebastian Pokutta , Zev Woodstock

We analyse a coupled 3D-2D model with a free fluid governed by Stokes flow in the bulk and a poroelastic plate described by the Biot-Kirchhoff equations on the surface. Assuming the form of a double perturbed saddle-point problem, the…

Numerical Analysis · Mathematics 2026-03-11 Franco Dassi , Rekha Khot , Andres E. Rubiano , Ricardo Ruiz-Baier

We introduce Monte-Carlo Attention (MCA), a randomized approximation method for reducing the computational cost of self-attention mechanisms in Transformer architectures. MCA exploits the fact that the importance of each token in an input…

Machine Learning · Computer Science 2022-02-01 Hyunjun Kim , JeongGil Ko

Maximal 't Hooft loops are studied in SO(3) lattice gauge theory at finite temperature T. Tunneling barriers among twist sectors causing loss of ergodicity for local update algorithms are overcome through parallel tempering, enabling us to…

High Energy Physics - Theory · Physics 2010-04-15 G. Burgio , M. Fuhrmann , W. Kerler , M. Muller-Preussker

We consider the dynamics of local entropy and nearest neighbor mutual information of a 1-D lattice of qubits via the repeated application of nearest neighbor CNOT quantum gates. This is a quantum version of a cellular automaton. We analyze…

Quantum Physics · Physics 2021-06-01 David Berenstein , Jiayao Zhao

Slow relaxation and glassiness have been the focus of extensive research attention, along with popular and technological interest, for many decades. While much understanding has been attained through mean-field and mode-coupling models,…

Disordered Systems and Neural Networks · Physics 2024-10-21 Alex Gower , Oliver Hart , Claudio Castelnovo

We investigate the thermodynamic and structural properties of divalent patchy hard rods confined to a one-dimensional channel by modeling the bonding sites as attractive square-well (SW) patches located at the rod tips. The zero-range…

Soft Condensed Matter · Physics 2026-05-21 Ana M. Montero , Andrés Santos , Péter Gurin , Szabolcs Varga

Despite the popularity of the Transformer architecture, the standard algorithm for computing Attention suffers from quadratic time complexity in context length $n$. Alman and Song [NeurIPS 2023] showed that when the head dimension $d =…

Machine Learning · Computer Science 2025-05-22 Shreya Gupta , Boyang Huang , Barna Saha , Yinzhan Xu , Christopher Ye

In a Hilbert setting, we introduce a new dynamical system and associated algorithms for solving monotone inclusions by rapid methods. Given a maximal monotone operator $A$, the evolution is governed by the time dependent operator $I -(I +…

Optimization and Control · Mathematics 2015-04-20 Hedy Attouch , Maicon Marques Alves , Benar F. Svaiter

Linear attention has attracted interest as a computationally efficient approximation to softmax attention, especially for long sequences. Recent studies have explored distilling softmax attention in pre-trained Transformers into linear…

Machine Learning · Computer Science 2025-07-08 Naoki Nishikawa , Rei Higuchi , Taiji Suzuki

In a Hilbert space setting H, for convex optimization, we analyze the fast convergence properties as t tends to infinity of the trajectories generated by a third-order in time evolution system. The function f to minimize is supposed to be…

Optimization and Control · Mathematics 2020-07-08 Hedy Attouch , Zaki Chbani , Hassan Riahi

Self-attention is a useful mechanism to build generative models for language and images. It determines the importance of context elements by comparing each element to the current time step. In this paper, we show that a very lightweight…

Computation and Language · Computer Science 2019-02-26 Felix Wu , Angela Fan , Alexei Baevski , Yann N. Dauphin , Michael Auli

We analyse a $2+1$ dimensional model with charged, relativistic fermions interacting through a four-Fermi term. Taking advantage of its large-$N$ renormalizability, the various phases of this model are studied at finite temperature and…

High Energy Physics - Theory · Physics 2010-11-01 R. MacKenzie , P. K. Panigrahi , S. Sakhi

Length-dependent logit rescaling is widely used to stabilize long-context self-attention, but existing analyses and methods suggest conflicting inverse-temperature laws for the context length $n$, ranging from $(\log n)^{1/2}$ to $\log n$…

Machine Learning · Statistics 2026-05-14 Tomohiro Hayase , Ryo Karakida