Related papers: Localmax dynamics for attention in transformers an…
The attention mechanism in Transformers is an important primitive for accurate and scalable sequence modeling. Its quadratic-compute and linear-memory complexity however remain significant bottlenecks. Linear attention and state-space…
We introduce a new class of nonlocal nonlinear conservation laws in one space dimension that allow for nonlocal interactions over a finite horizon. The proposed model, which we refer to as the nonlocal pair interaction model, inherits at…
Fixed-time stable dynamical systems are capable of achieving exact convergence to an equilibrium point within a fixed time that is independent of the initial conditions of the system. This property makes them highly appealing for designing…
We study a family of local depth-based corrections to maxmin landmark selection for lazy witness persistence. Starting from maxmin seeds, we partition the cloud into nearest-seed cells and replace or move each seed toward a deep…
We investigate the late-time asymptotic behavior of solutions to nonlinear hyperbolic systems of conservation laws containing stiff relaxation terms. First, we introduce a Chapman-Enskog-type asymptotic expansion and derive an effective…
Transformers have achieved great success in recent years. Interestingly, transformers have shown particularly strong in-context learning capability -- even without fine-tuning, they are still able to solve unseen tasks well purely based on…
We study equilibrium selection for invariant measures of stochastic dynamical systems with constant step size, under persistent noise and minimal moment assumptions, in a general quasi-Feller framework. Such dynamics arise in…
Recent developments in topological mechanics have demonstrated the ability of Maxwell lattices to effectively focus stress along domain walls between differently polarized domains. The focusing ability can be exploited to protect the…
We study the training dynamics of gradient descent in a softmax self-attention layer trained to perform linear regression and show that a simple first-order optimization algorithm can converge to the globally optimal self-attention…
Attention is a key part of the transformer architecture. It is a sequence-to-sequence mapping that transforms each sequence element into a weighted sum of values. The weights are typically obtained as the softmax of dot products between…
This paper presents $\textbf{CAPS}$ (Clock-weighted Aggregation with Prefix-products and Softmax), a structured attention mechanism for time series forecasting that decouples three distinct temporal structures: global trends, local shocks,…
This study presents a constructive methodology for designing accelerated convex optimisation algorithms in continuous-time domain. The two key enablers are the classical concept of passivity in control theory and the time-dependent change…
New sufficient conditions for the characterization of dwell-times for linear impulsive systems are proposed and shown to coincide with continuous decrease conditions of a certain class of looped-functionals, a recently introduced type of…
The Transformer architecture, a cornerstone of modern Large Language Models (LLMs), has achieved extraordinary success in sequence modeling, primarily due to its attention mechanism. However, despite its power, the standard attention…
Transformers are state-of-the-art in a wide range of NLP tasks and have also been applied to many real-world products. Understanding the reliability and certainty of transformer model predictions is crucial for building trustable machine…
We study the stability of one-dimensional linear hyperbolic systems with non-symmetric relaxation. Introducing a new frequency-dependent Kalman stability condition, we prove an abstract decay result underpinning a form of inhomogeneous…
Adaptive control architectures often make use of Lyapunov functions to design adaptive laws. We are specifically interested in adaptive control methods, such as the well-known L1 adaptive architecture, which employ a parameter observer for…
To address the communication bottleneck problem in distributed optimization within a master-worker framework, we propose LocalNewton, a distributed second-order algorithm with local averaging. In LocalNewton, the worker machines update…
In this paper, we extend the standard Attention in transformer by exploiting the consensus discrepancy from a distributed optimization perspective, referred to as AttentionX. It is noted that the primal-dual method of multipliers (PDMM)…
Softmax feedback systems are a common mathematical core of entropy-regularized reinforcement learning, logit game dynamics, population choice, and mean-field variational updates. Their central stability question is simple: when does a…