English
Related papers

Related papers: Uniform Scaling Limits in AdamW-Trained Transforme…

200 papers

Time series alignment methods call for highly expressive, differentiable and invertible warping functions which preserve temporal topology, i.e diffeomorphisms. Diffeomorphic warping functions can be generated from the integration of…

Machine Learning · Computer Science 2022-06-17 Iñigo Martinez , Elisabeth Viles , Igor G. Olaizola

In this paper, we develop a novel argument, the non-autonomous approximation method, to seek the asymptotic limits of the fully coupled multi-scale McKean-Vlasov stochastic systems with irregular coefficients, which, as summarized in…

Probability · Mathematics 2024-12-19 Yuewen Hou , Yun Li , Longjie Xie

In this paper we delve deep in the Transformer architecture by investigating two of its core components: self-attention and contextual embeddings. In particular, we study the identifiability of attention weights and token embeddings, and…

Computation and Language · Computer Science 2020-02-10 Gino Brunner , Yang Liu , Damián Pascual , Oliver Richter , Massimiliano Ciaramita , Roger Wattenhofer

In a recent paper by Guglielmi and Hairer (SIADS 2015), an analysis in the $\varepsilon\to 0$ limit was proposed of regularized discontinuous ODEs in codimension-2 switching domains; this was obtained by studying a certain 2-dimensional…

Dynamical Systems · Mathematics 2024-06-04 Alessia andò , Roderick Edwards , Nicola Guglielmi

In recent years, increasingly large models have achieved outstanding performance across CV tasks. However, these models demand substantial computational resources and storage, and their growing complexity limits our understanding of how…

Machine Learning · Computer Science 2025-11-21 Carlos Boned Riera , David Romero Sanchez , Oriol Ramos Terrades

Stochastic differential equations (SDE) often exhibit large random transitions. This property, which we denote as pathwise stiffness, causes transient bursts of stiffness which limit the allowed step size for common fixed time step explicit…

Numerical Analysis · Mathematics 2018-04-13 Christopher Rackauckas , Qing Nie

We develop a mathematical framework that interprets Transformer attention as an interacting particle system and studies its continuum (mean-field) limits. By idealizing attention on the sphere, we connect Transformer dynamics to Wasserstein…

Machine Learning · Computer Science 2026-02-02 Philippe Rigollet

Compressed Deep Learning (DL) models are essential for deployment in resource-constrained environments. But their performance often lags behind their large-scale counterparts. To bridge this gap, we propose Alignment Adapter (AlAd): a…

Machine Learning · Computer Science 2026-02-17 Rohit Raj Rai , Abhishek Dhaka , Amit Awekar

The multi-head attention layer is one of the key components of the transformer architecture that sets it apart from traditional feed-forward models. Given a sequence length $k$, attention matrices…

Machine Learning · Computer Science 2024-02-07 Sitan Chen , Yuanzhi Li

Amodal depth estimation aims to predict the depth of occluded (invisible) parts of objects in a scene. This task addresses the question of whether models can effectively perceive the geometry of occluded regions based on visible cues. Prior…

Computer Vision and Pattern Recognition · Computer Science 2024-12-04 Zhenyu Li , Mykola Lavreniuk , Jian Shi , Shariq Farooq Bhat , Peter Wonka

We investigate two one-dimensional tight-binding models with disorder that have extended states at zero energy. We use exact and partial diagonalisation of the Hamiltonian to obtain the eigenmodes and the associated participation ratios,…

Disordered Systems and Neural Networks · Physics 2025-08-27 Luca Schaefer , Barbara Drossel

Neural operators offer a powerful data-driven framework for learning mappings between function spaces, in which the transformer-based neural operator architecture faces a fundamental scalability-accuracy trade-off: softmax attention…

Machine Learning · Computer Science 2025-10-21 Ming Zhong , Zhenya Yan

Augmenting mechanistic ordinary differential equation (ODE) models with machine-learnable structures is an novel approach to create highly accurate, low-dimensional models of engineering systems incorporating both expert knowledge and…

Dynamical Systems · Mathematics 2022-06-22 Sandor Beregi , David A. W. Barton , Djamel Rezgui , Simon A. Neild

Despite the promise of scientific machine learning (SciML) in combining data-driven techniques with mechanistic modeling, existing approaches for incorporating hard constraints in neural differential equations (NDEs) face significant…

Machine Learning · Computer Science 2025-05-28 Avik Pal , Alan Edelman , Christopher Rackauckas

Many important problems in science and engineering require solving the so-called parametric partial differential equations (PDEs), i.e., PDEs with different physical parameters, boundary conditions, shapes of computational domains, etc.…

Numerical Analysis · Mathematics 2024-02-06 Zhanhong Ye , Xiang Huang , Hongsheng Liu , Bin Dong

Training of discrete latent variable models remains challenging because passing gradient information through discrete units is difficult. We propose a new class of smoothing transformations based on a mixture of two overlapping…

Machine Learning · Computer Science 2018-05-29 Arash Vahdat , William G. Macready , Zhengbing Bian , Amir Khoshaman , Evgeny Andriyash

As emerging quantum architectures evolve into heterogeneous networks combining different physical substrates, such as qubits for logic and higher-dimensional qudits for robust communication, the traditional scalar metrics of quantum error…

Quantum Physics · Physics 2026-04-29 David González-Lociga , Simeon Ball

The canonical $O(N^2)$ Transformer remains the empirical performance frontier in sequence modeling, and its training can be further optimized by addressing geometric inefficiency. We propose an optimization framework that leverages an…

Machine Learning · Computer Science 2025-12-16 Jongwook Kim , Sangheon Yun , Sukjin Yoon

At low temperature T, a significant difference between the behavior of crystals on the one hand and disordered solids on the other is seen: sufficiently strong disorder can give rise to a transition of the transport properties from…

Disordered Systems and Neural Networks · Physics 2018-03-21 Rudolf A Roemer , Michael Schreiber

Modern large language models (LLMs) excel at tasks that require storing and retrieving knowledge, such as factual recall and question answering. Transformers are central to this capability because they can encode information during training…

Machine Learning · Statistics 2026-03-18 Nuri Mert Vural , Alberto Bietti , Mahdi Soltanolkotabi , Denny Wu