English
Related papers

Related papers: Localmax dynamics for attention in transformers an…

200 papers

The attention mechanism in Transformers is an important primitive for accurate and scalable sequence modeling. Its quadratic-compute and linear-memory complexity however remain significant bottlenecks. Linear attention and state-space…

Machine Learning · Computer Science 2026-03-03 Han Guo , Songlin Yang , Tarushii Goel , Eric P. Xing , Tri Dao , Yoon Kim

We introduce a new class of nonlocal nonlinear conservation laws in one space dimension that allow for nonlocal interactions over a finite horizon. The proposed model, which we refer to as the nonlocal pair interaction model, inherits at…

Analysis of PDEs · Mathematics 2016-11-29 Qiang Du , Zhan Huang , Philippe G. LeFloch

Fixed-time stable dynamical systems are capable of achieving exact convergence to an equilibrium point within a fixed time that is independent of the initial conditions of the system. This property makes them highly appealing for designing…

Systems and Control · Electrical Eng. & Systems 2025-10-01 Michael Tang , Miroslav Krstic , Jorge Poveda

We study a family of local depth-based corrections to maxmin landmark selection for lazy witness persistence. Starting from maxmin seeds, we partition the cloud into nearest-seed cells and replace or move each seed toward a deep…

Computational Geometry · Computer Science 2026-04-22 Yifan Zhang

We investigate the late-time asymptotic behavior of solutions to nonlinear hyperbolic systems of conservation laws containing stiff relaxation terms. First, we introduce a Chapman-Enskog-type asymptotic expansion and derive an effective…

Analysis of PDEs · Mathematics 2011-09-20 Christophe Berthon , Philippe G. LeFloch , Rodolphe Turpault

Transformers have achieved great success in recent years. Interestingly, transformers have shown particularly strong in-context learning capability -- even without fine-tuning, they are still able to solve unseen tasks well purely based on…

Machine Learning · Computer Science 2024-11-19 Zihao Li , Yuan Cao , Cheng Gao , Yihan He , Han Liu , Jason M. Klusowski , Jianqing Fan , Mengdi Wang

We study equilibrium selection for invariant measures of stochastic dynamical systems with constant step size, under persistent noise and minimal moment assumptions, in a general quasi-Feller framework. Such dynamics arise in…

Probability · Mathematics 2026-01-16 Jean-Gabriel Attali

Recent developments in topological mechanics have demonstrated the ability of Maxwell lattices to effectively focus stress along domain walls between differently polarized domains. The focusing ability can be exploited to protect the…

Soft Condensed Matter · Physics 2025-02-04 Caleb Widstrand , Xiaoming Mao , Stefano Gonella

We study the training dynamics of gradient descent in a softmax self-attention layer trained to perform linear regression and show that a simple first-order optimization algorithm can converge to the globally optimal self-attention…

Machine Learning · Computer Science 2026-03-03 Gautam Goel , Mahdi Soltanolkotabi , Peter Bartlett

Attention is a key part of the transformer architecture. It is a sequence-to-sequence mapping that transforms each sequence element into a weighted sum of values. The weights are typically obtained as the softmax of dot products between…

This paper presents $\textbf{CAPS}$ (Clock-weighted Aggregation with Prefix-products and Softmax), a structured attention mechanism for time series forecasting that decouples three distinct temporal structures: global trends, local shocks,…

Machine Learning · Computer Science 2026-02-04 Viresh Pati , Yubin Kim , Vinh Pham , Jevon Twitty , Shihao Yang , Jiecheng Lu

This study presents a constructive methodology for designing accelerated convex optimisation algorithms in continuous-time domain. The two key enablers are the classical concept of passivity in control theory and the time-dependent change…

Optimization and Control · Mathematics 2024-09-16 Namhoon Cho , Hyo-Sang Shin

New sufficient conditions for the characterization of dwell-times for linear impulsive systems are proposed and shown to coincide with continuous decrease conditions of a certain class of looped-functionals, a recently introduced type of…

Optimization and Control · Mathematics 2012-06-05 Corentin Briat , Alexandre Seuret

The Transformer architecture, a cornerstone of modern Large Language Models (LLMs), has achieved extraordinary success in sequence modeling, primarily due to its attention mechanism. However, despite its power, the standard attention…

Machine Learning · Computer Science 2026-01-08 Zichuan Fu , Wentao Song , Guojing Li , Yejing Wang , Xian Wu , Yimin Deng , Hanyu Yan , Yefeng Zheng , Xiangyu Zhao

Transformers are state-of-the-art in a wide range of NLP tasks and have also been applied to many real-world products. Understanding the reliability and certainty of transformer model predictions is crucial for building trustable machine…

Computation and Language · Computer Science 2021-12-28 Jiahuan Pei , Cheng Wang , György Szarvas

We study the stability of one-dimensional linear hyperbolic systems with non-symmetric relaxation. Introducing a new frequency-dependent Kalman stability condition, we prove an abstract decay result underpinning a form of inhomogeneous…

Analysis of PDEs · Mathematics 2025-03-04 Timothée Crin-Barat , Lorenzo Liverani , Ling-Yun Shou , Enrique Zuazua

Adaptive control architectures often make use of Lyapunov functions to design adaptive laws. We are specifically interested in adaptive control methods, such as the well-known L1 adaptive architecture, which employ a parameter observer for…

Systems and Control · Electrical Eng. & Systems 2021-04-23 Aditya A. Paranjape , Vivek Natarajan , Supratim Ghosh

To address the communication bottleneck problem in distributed optimization within a master-worker framework, we propose LocalNewton, a distributed second-order algorithm with local averaging. In LocalNewton, the worker machines update…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-05-18 Vipul Gupta , Avishek Ghosh , Michal Derezinski , Rajiv Khanna , Kannan Ramchandran , Michael Mahoney

In this paper, we extend the standard Attention in transformer by exploiting the consensus discrepancy from a distributed optimization perspective, referred to as AttentionX. It is noted that the primal-dual method of multipliers (PDMM)…

Machine Learning · Computer Science 2024-10-15 Guoqiang Zhang , Richard Heusdens

Softmax feedback systems are a common mathematical core of entropy-regularized reinforcement learning, logit game dynamics, population choice, and mean-field variational updates. Their central stability question is simple: when does a…

Machine Learning · Computer Science 2026-05-18 Tongxi Wang