English
Related papers

Related papers: Uniform Scaling Limits in AdamW-Trained Transforme…

200 papers

Despite the vast empirical evidence supporting the efficacy of adaptive optimization methods in deep learning, their theoretical understanding is far from complete. This work introduces novel SDEs for commonly used adaptive optimizers:…

Machine Learning · Computer Science 2025-03-12 Enea Monzio Compagnoni , Tianlin Liu , Rustem Islamov , Frank Norbert Proske , Antonio Orvieto , Aurelien Lucchi

Adaptive control of Euler-Lagrange systems is challenging when friction is governed by a finite-horizon internal state that is not directly observable from joint measurements. In this setting, the measured closed-loop state is no longer…

Machine Learning · Computer Science 2026-05-11 Giansalvo Cirrincione , Adriano Fagiolini

We present the Condensate Theorem: attention sparsity is a learned topological property, not an architectural constraint. Through empirical analysis of trained language models, we find that attention mass concentrates on a distinct…

Machine Learning · Computer Science 2026-02-11 Jorge L. Ruiz Williams

The Bose-Hubbard model (BHM) has been widely explored to develop a profound understanding of the strongly correlated behavior of interacting bosons. Quantum simulators not only allow the exploration of the BHM but also extend it to models…

Quantum Gases · Physics 2024-03-12 Zeki Zeybek , Peter Schmelcher , Rick Mukherjee

In designing efficient feedback control laws for fluid flow, the modern control theory can serve as a powerful tool if the model can be represented by a linear ordinary differential equation (ODE). However, it is generally difficult to find…

Fluid Dynamics · Physics 2023-11-16 Hiroshi Omichi , Takeru Ishize , Koji Fukagata

An oblivious subspace embedding (OSE), characterized by parameters $m,n,d,\epsilon,\delta$, is a random matrix $\Pi\in \mathbb{R}^{m\times n}$ such that for any $d$-dimensional subspace $T\subseteq \mathbb{R}^n$, $\Pr_\Pi[\forall x\in T,…

Data Structures and Algorithms · Computer Science 2023-07-14 Yi Li , Mingmou Liu

We study two strange phenomena in auto-regressive Transformers: (1) the dominance of the first token in attention heads; (2) the occurrence of large outlier activations in the hidden states. We find that popular large language models, such…

Computation and Language · Computer Science 2024-10-23 Prannay Kaul , Chengcheng Ma , Ismail Elezi , Jiankang Deng

We study the robustness of Transformer language models under semantic out-of-distribution (OOD) shifts, where training and test data lie in disjoint latent spaces. Using Wasserstein-1 distance and Gevrey-class smoothness, we derive…

Machine Learning · Computer Science 2025-06-02 Yu Wang , Fu-Chieh Chang , Pei-Yuan Wu

The rigorous linking of exact stochastic models to mean-field approximations is studied. Starting from the differential equation point of view the stochastic model is identified by its Kolmogorov equations, which is a system of linear ODEs…

Dynamical Systems · Mathematics 2011-09-19 András Bátkai , Istvan Z. Kiss , Eszter Sikolya , Péter L. Simon

The Adam optimizer is currently presumably the most popular optimization method in deep learning. In this article we develop an ODE based method to study the Adam optimizer in a fast-slow scaling regime. For fixed momentum parameters and…

Optimization and Control · Mathematics 2025-11-07 Steffen Dereich , Arnulf Jentzen , Sebastian Kassing

Transformer models have become the dominant backbone for sequence modeling, leveraging self-attention to produce contextualized token representations. These are typically aggregated into fixed-size vectors via pooling operations for…

Machine Learning · Computer Science 2025-10-07 Sofiane Ennadir , Levente Zólyomi , Oleg Smirnov , Tianze Wang , John Pertoft , Filip Cornell , Lele Cao

We study causal self-attention dynamics -- a toy model for decoder Transformers -- which we interpret as a non-exchangeable interacting particle system. Adapting cumulant expansions to the triangular causal dependency structure of the…

Analysis of PDEs · Mathematics 2026-05-12 Mitia Duerinckx , Borjan Geshkovski , Stefano Rossi

We present an exact dimensional reduction for high-dimensional dynamical systems composed of $N$ identical dynamical units governed by quasi-linear ordinary differential equations (ODEs) of order $M$. In these systems, each unit follows a…

Adaptation and Self-Organizing Systems · Physics 2026-05-15 Felix Augustsson , Erik Andreas Martens , Rok Cestnik

Transformer-based models have achieved remarkable success across a wide range of domains, yet our understanding of their training dynamics remains limited. In this work, we identify a recurrent focus-dilution cycle in attention learning and…

Machine Learning · Computer Science 2026-05-05 Zheng-An Chen , Pengxiao Lin , Zhi-Qin John Xu , Tao Luo

We study the quantum (zero-temperature) critical behaviors of confined particle systems described by the one-dimensional (1D) Bose-Hubbard model in the presence of a confining potential, at the Mott insulator to superfluid transitions, and…

Statistical Mechanics · Physics 2013-05-29 Massimo Campostrini , Ettore Vicari

We study McKean--Vlasov Stochastic Differential Equations (MV-SDEs) whose drift and diffusion coefficients are of superlinear growth in \textit{all} their variables thus also superlinear in the measure component (the meaning is specified in…

Probability · Mathematics 2025-10-21 Simran Soni , Neelima , Chaman Kumar , Goncalo dos Reis

This paper studies the contraction property of time-varying differential-algebraic equation (DAE) systems by embedding them to higher-dimension ordinary differential equation (ODE) systems. The first result pertains to the equivalence of…

Systems and Control · Electrical Eng. & Systems 2023-12-12 Hao Yin , Bayu Jayawardhana , Stephan Trenn

Transformer-based models are becoming a central paradigm in autonomous driving because they can capture long-range spatial dependencies, multi-agent interactions, and multimodal context across perception, prediction, and planning. At the…

Machine Learning · Computer Science 2026-05-13 Juan Zhong , Yuhang Shi , Zukang Xu , Xi Chen

With the rise of Transformer models in NLP and CV domain, Multi-Head Attention has been proven to be a game-changer. However, its expensive computation poses challenges to the model throughput and efficiency, especially for the long…

Image and Video Processing · Electrical Eng. & Systems 2024-04-12 Jiing-Ping Wang , Ming-Guang Lin , An-Yeu , Wu

We develop a reduced-order framework for optimizing mixing in two-dimensional incompressible flows. Instead of optimizing the full transport PDE, the method maximizes the length of advected material interfaces, leading to a…

Numerical Analysis · Mathematics 2026-05-07 Ziqian Li , Enrique Zuazua
‹ Prev 1 4 5 6 7 8 10 Next ›