中文
相关论文

相关论文: Attention's forward pass and Frank-Wolfe

200 篇论文

Large Language Models (LLMs) exhibit exceptional proficiency in handling extensive context windows in natural language. Nevertheless, the quadratic scaling of attention computation relative to sequence length creates substantial efficiency…

机器学习 · 计算机科学 2026-01-26 Xiaoyu Li , Yingyu Liang , Zhenmei Shi , Zhao Song , Song Yue , Jiahao Zhang

Transformers are state-of-the-art in a wide range of NLP tasks and have also been applied to many real-world products. Understanding the reliability and certainty of transformer model predictions is crucial for building trustable machine…

计算与语言 · 计算机科学 2021-12-28 Jiahuan Pei , Cheng Wang , György Szarvas

We develop a new numerical method for thin plates falling in inviscid fluid that allows for leading-edge vortex shedding. The inclusion of leading-edge shedding restores physical dynamics to vortex-sheet models of falling bodies, and for…

流体动力学 · 物理学 2025-10-02 Yu Jun Loo , Silas Alben

Transformer-based models have achieved remarkable success across a wide range of domains, yet our understanding of their training dynamics remains limited. In this work, we identify a recurrent focus-dilution cycle in attention learning and…

机器学习 · 计算机科学 2026-05-05 Zheng-An Chen , Pengxiao Lin , Zhi-Qin John Xu , Tao Luo

Tensor Attention extends traditional attention mechanisms by capturing high-order correlations across multiple modalities, addressing the limitations of classical matrix-based attention. Meanwhile, Rotary Position Embedding…

机器学习 · 计算机科学 2024-12-25 Xiaoyu Li , Yingyu Liang , Zhenmei Shi , Zhao Song , Mingda Wan

We develop an algorithmic framework for solving convex optimization problems using no-regret game dynamics. By converting the problem of minimizing a convex function into an auxiliary problem of solving a min-max game in a sequential…

机器学习 · 计算机科学 2023-02-21 Jun-Kun Wang , Jacob Abernethy , Kfir Y. Levy

We prove that the block-coordinate Frank-Wolfe (BCFW) algorithm converges with state-of-the-art rates in both convex and nonconvex settings under a very mild "block-iterative" assumption. This appears to be the first result on BCFW…

最优化与控制 · 数学 2025-12-17 Gábor Braun , Jannis Halbey , Sebastian Pokutta , Zev Woodstock

We analyse a coupled 3D-2D model with a free fluid governed by Stokes flow in the bulk and a poroelastic plate described by the Biot-Kirchhoff equations on the surface. Assuming the form of a double perturbed saddle-point problem, the…

数值分析 · 数学 2026-03-11 Franco Dassi , Rekha Khot , Andres E. Rubiano , Ricardo Ruiz-Baier

We introduce Monte-Carlo Attention (MCA), a randomized approximation method for reducing the computational cost of self-attention mechanisms in Transformer architectures. MCA exploits the fact that the importance of each token in an input…

机器学习 · 计算机科学 2022-02-01 Hyunjun Kim , JeongGil Ko

Maximal 't Hooft loops are studied in SO(3) lattice gauge theory at finite temperature T. Tunneling barriers among twist sectors causing loss of ergodicity for local update algorithms are overcome through parallel tempering, enabling us to…

高能物理 - 理论 · 物理学 2010-04-15 G. Burgio , M. Fuhrmann , W. Kerler , M. Muller-Preussker

We consider the dynamics of local entropy and nearest neighbor mutual information of a 1-D lattice of qubits via the repeated application of nearest neighbor CNOT quantum gates. This is a quantum version of a cellular automaton. We analyze…

量子物理 · 物理学 2021-06-01 David Berenstein , Jiayao Zhao

Slow relaxation and glassiness have been the focus of extensive research attention, along with popular and technological interest, for many decades. While much understanding has been attained through mean-field and mode-coupling models,…

无序系统与神经网络 · 物理学 2024-10-21 Alex Gower , Oliver Hart , Claudio Castelnovo

We investigate the thermodynamic and structural properties of divalent patchy hard rods confined to a one-dimensional channel by modeling the bonding sites as attractive square-well (SW) patches located at the rod tips. The zero-range…

软凝聚态物质 · 物理学 2026-05-21 Ana M. Montero , Andrés Santos , Péter Gurin , Szabolcs Varga

Despite the popularity of the Transformer architecture, the standard algorithm for computing Attention suffers from quadratic time complexity in context length $n$. Alman and Song [NeurIPS 2023] showed that when the head dimension $d =…

机器学习 · 计算机科学 2025-05-22 Shreya Gupta , Boyang Huang , Barna Saha , Yinzhan Xu , Christopher Ye

In a Hilbert setting, we introduce a new dynamical system and associated algorithms for solving monotone inclusions by rapid methods. Given a maximal monotone operator $A$, the evolution is governed by the time dependent operator $I -(I +…

最优化与控制 · 数学 2015-04-20 Hedy Attouch , Maicon Marques Alves , Benar F. Svaiter

Linear attention has attracted interest as a computationally efficient approximation to softmax attention, especially for long sequences. Recent studies have explored distilling softmax attention in pre-trained Transformers into linear…

机器学习 · 计算机科学 2025-07-08 Naoki Nishikawa , Rei Higuchi , Taiji Suzuki

In a Hilbert space setting H, for convex optimization, we analyze the fast convergence properties as t tends to infinity of the trajectories generated by a third-order in time evolution system. The function f to minimize is supposed to be…

最优化与控制 · 数学 2020-07-08 Hedy Attouch , Zaki Chbani , Hassan Riahi

Self-attention is a useful mechanism to build generative models for language and images. It determines the importance of context elements by comparing each element to the current time step. In this paper, we show that a very lightweight…

计算与语言 · 计算机科学 2019-02-26 Felix Wu , Angela Fan , Alexei Baevski , Yann N. Dauphin , Michael Auli

We analyse a $2+1$ dimensional model with charged, relativistic fermions interacting through a four-Fermi term. Taking advantage of its large-$N$ renormalizability, the various phases of this model are studied at finite temperature and…

高能物理 - 理论 · 物理学 2010-11-01 R. MacKenzie , P. K. Panigrahi , S. Sakhi

Length-dependent logit rescaling is widely used to stabilize long-context self-attention, but existing analyses and methods suggest conflicting inverse-temperature laws for the context length $n$, ranging from $(\log n)^{1/2}$ to $\log n$…

机器学习 · 统计学 2026-05-14 Tomohiro Hayase , Ryo Karakida