English
Related papers

Related papers: Attention in Krylov Space

200 papers

I show that ordinary least squares (OLS) predictions can be rewritten as the output of a restricted attention module, akin to those forming the backbone of large language models. This connection offers an alternative perspective on…

Machine Learning · Computer Science 2026-01-13 Philippe Goulet Coulombe

Data augmentation, a technique in which a training set is expanded with class-preserving transformations, is ubiquitous in modern machine learning pipelines. In this paper, we seek to establish a theoretical framework for understanding data…

Machine Learning · Computer Science 2019-03-21 Tri Dao , Albert Gu , Alexander J. Ratner , Virginia Smith , Christopher De Sa , Christopher Ré

Attention patterns play a crucial role in both training and inference of large language models (LLMs). Prior works have identified individual patterns such as retrieval heads, sink heads, and diagonal traces, yet these observations remain…

Computation and Language · Computer Science 2026-01-30 Qingyue Yang , Jie Wang , Xing Li , Yinqi Bai , Xialiang Tong , Huiling Zhen , Jianye Hao , Mingxuan Yuan , Bin Li

Transformers are increasingly adopted for modeling and forecasting time-series, yet their internal mechanisms remain poorly understood from a dynamical systems perspective. In contrast to classical autoregressive and state-space models,…

Machine Learning · Computer Science 2025-12-25 Gregory Duthé , Nikolaos Evangelou , Wei Liu , Ioannis G. Kevrekidis , Eleni Chatzi

F$^{19}$ nuclear magnetic resonance free induction decay (FID) data are used to verify the predictions of a universal growth hypothesis for the Lanczos coefficients proposed by Parker et al. Our results strongly support this hypothesis and…

Mesoscale and Nanoscale Physics · Physics 2026-04-13 M. Engelsberg , Wilson Barros

We develop a deterministic large-time mechanism yielding Ces{\`a}ro asymptotic observability inequalities from moving localized observations for conservative evolutions. On each observation interval, exact convexification on a compact…

Analysis of PDEs · Mathematics 2026-05-13 Maarten V. de Hoop , Antti Kykkänen , Emmanuel Trélat

While attention has been empirically shown to improve model performance, it lacks a rigorous mathematical justification. This short paper establishes a novel connection between attention mechanisms and multinomial regression. Specifically,…

Machine Learning · Computer Science 2025-10-28 Jonas A. Actor , Anthony Gruber , Eric C. Cyr

The time-ordered exponential of a time-dependent matrix $\mathsf{A}(t)$ is defined as the function of $\mathsf{A}(t)$ that solves the first-order system of coupled linear differential equations with non-constant coefficients encoded in…

Numerical Analysis · Mathematics 2020-10-09 Pierre-Louis Giscard , Stefano Pozza

Learning reduced descriptions of chaotic many-body dynamics is fundamentally challenging: although microscopic equations are Markovian, collective observables exhibit strong memory and exponential sensitivity to initial conditions and…

Computational Physics · Physics 2026-01-28 Ho Jang , Gia-Wei Chern

Compared to the classical Lanczos algorithm, the $s$-step Lanczos variant has the potential to improve performance by asymptotically decreasing the synchronization cost per iteration. However, this comes at a cost. Despite being…

Numerical Analysis · Mathematics 2021-08-31 Erin Carson , Tomáš Gergelits

Fourier Neural Operators (FNOs) excel on tasks using functional data, such as those originating from partial differential equations. Such characteristics render them an effective approach for simulating the time evolution of quantum…

In this paper, we present a new approach for model reduction of large scale first and second order dynamical systems with multiple inputs and multiple outputs (MIMO). This approach is based on the projection of the initial problem onto…

Numerical Analysis · Computer Science 2019-03-19 Yassine Kaouane , Khalide Jbilou

Continual learning aims to sequentially learn new tasks without forgetting previous tasks' knowledge (catastrophic forgetting). One factor that can cause forgetting is the interference between the gradients on losses from different tasks.…

Computation and Language · Computer Science 2025-12-01 Xueying Bai , Jinghuan Shang , Yifan Sun , Niranjan Balasubramanian

Recursive stochastic algorithms have gained significant attention in the recent past due to data driven applications. Examples include stochastic gradient descent for solving large-scale optimization problems and empirical dynamic…

Machine Learning · Computer Science 2020-07-27 Abhishek Gupta , Hao Chen , Jianzong Pi , Gaurav Tendolkar

In this paper, we study a class of convolution operators on the space of distributions that enlarge the well-studied class of passive operators. In this larger class, we are able to associate, to each operator, a holomorphic function in the…

Functional Analysis · Mathematics 2018-11-27 Mitja Nedic

Thermalization and scrambling are the subject of much recent study from the perspective of many-body quantum systems with locally bounded Hilbert spaces (`spin chains'), quantum field theory and holography. We tackle this problem in 1D…

Strongly Correlated Electrons · Physics 2018-04-18 Curt von Keyserlingk , Tibor Rakovszky , Frank Pollmann , Shivaji Sondhi

Reliable adaptive beamforming is critical for large microphone arrays operating in highly dynamic acoustic environments. In scenarios characterized by fast-moving talkers and interferers, the available sample support for estimating the…

Signal Processing · Electrical Eng. & Systems 2026-05-13 Manan Mittal , Ryan M. Corey , John R. Buck , Andrew C. Singer

We propose a method for learning dynamical systems from high-dimensional empirical data that combines variational autoencoders and (spatio-)temporal attention within a framework designed to enforce certain scientifically-motivated…

Machine Learning · Computer Science 2023-06-22 Kai Lagemann , Christian Lagemann , Sach Mukherjee

We formulate an attention mechanism for continuous and ordered sequences that explicitly functions as an alignment model, which serves as the core of many sequence-to-sequence tasks. Standard scaled dot-product attention relies on…

Machine Learning · Computer Science 2025-09-19 Hyungjoon Soh , Junghyo Jo

We propose a synthetic reasoning task, LEGO (Learning Equality and Group Operations), that encapsulates the problem of following a chain of reasoning, and we study how the Transformer architectures learn this task. We pay special attention…

Machine Learning · Computer Science 2023-02-21 Yi Zhang , Arturs Backurs , Sébastien Bubeck , Ronen Eldan , Suriya Gunasekar , Tal Wagner