English
Related papers

Related papers: Uniform Scaling Limits in AdamW-Trained Transforme…

200 papers

Transformers, which are state-of-the-art in most machine learning tasks, represent the data as sequences of vectors called tokens. This representation is then exploited by the attention function, which learns dependencies between tokens and…

Machine Learning · Computer Science 2025-01-31 Valérie Castin , Pierre Ablin , José Antonio Carrillo , Gabriel Peyré

Recent advancements in large language models (LLMs) based on transformer architectures have sparked significant interest in understanding their inner workings. In this paper, we introduce a novel approach to modeling transformer…

Machine Learning · Computer Science 2025-04-17 Anh Tong , Thanh Nguyen-Tang , Dongeun Lee , Duc Nguyen , Toan Tran , David Hall , Cheongwoong Kang , Jaesik Choi

Adaptive optimizers like AdamW apply uniform hyperparameters across all parameter groups, ignoring heterogeneous optimization dynamics across layers and modules. We address this limitation by proposing MetaAdamW - a new optimizer that…

Machine Learning · Computer Science 2026-05-07 JiangBo Zhao , ZhaoXin Liu

We present a hybrid transformer architecture that replaces discrete middle layers with a continuous-depth Neural Ordinary Differential Equation (ODE) block, enabling inference-time control over generation attributes via a learned steering…

Machine Learning · Computer Science 2026-01-16 Peter Jemley

We study a random model of deep multi-head self-attention in which the weights are resampled independently across layers and heads, as at initialization of training. Viewing depth as a time variable, the residual stream defines a…

Probability · Mathematics 2026-04-03 Hugo Koubbi , Borjan Geshkovski , Philippe Rigollet

Foundation models have transformed language, vision, and time series data analysis, yet progress on dynamic predictions for physical systems remains limited. Given the complexity of physical constraints, two challenges stand out. $(i)$…

Machine Learning · Computer Science 2026-02-05 Haoran Li , Chenhan Xiao , Lihao Mai , Yang Weng , Erik Blasch

In deep learning theory, the covariance matrix of the representations serves as a proxy to examine the network's trainability. Motivated by the success of Transformers, we study the covariance matrix of a modified Softmax-based attention…

Machine Learning · Statistics 2023-12-12 Lorenzo Noci , Chuning Li , Mufan Bill Li , Bobby He , Thomas Hofmann , Chris Maddison , Daniel M. Roy

Transformers excel at time series modelling through attention mechanisms that capture long-term temporal patterns. However, they assume uniform time intervals and therefore struggle with irregular time series. Neural Ordinary Differential…

Machine Learning · Computer Science 2026-05-13 Yashas Shende , Aritra Das , Reva Laxmi Chauhan , Arghya Pathak , Debayan Gupta

We analyze cumulative parameter trajectories of transformer training under AdamW and identify a dominant low-dimensional drift direction ("backbone") that captures 60--80% of long-horizon displacement from initialization. This direction is…

Machine Learning · Computer Science 2026-03-20 Yongzhong Xu

Transformers have become the dominant architecture in modern machine learning, yet the theoretical understanding of their training dynamics remains limited. This paper develops a rigorous mathematical framework for analyzing gradient-based…

Optimization and Control · Mathematics 2026-05-19 Raphaël Barboni , Maarten V. de Hoop , Takashi Furuya , Gabriel Peyré

In this paper, we develop a rigorous optimal control-theoretic approach to Transformer training that respects key structural constraints such as (i) realized-input-independence during execution, (ii) the ensemble control nature of the…

Machine Learning · Computer Science 2026-03-11 Kağan Akman , Naci Saldı , Serdar Yüksel

This letter presents a high-dimensional analysis of the training dynamics for a single-layer nonlinear contrastive learning model. The empirical distribution of the model weights converges to a deterministic measure governed by a…

Machine Learning · Computer Science 2024-06-12 Lineghuan Meng , Chuang Wang

We investigate the asymptotic properties of deep Residual networks (ResNets) as the number of layers increases. We first show the existence of scaling regimes for trained weights markedly different from those implicitly assumed in the…

Machine Learning · Computer Science 2023-01-26 Rama Cont , Alain Rossier , Renyuan Xu

We develop a backstepping-based observer design for a class of ODE - continuum-PDE cascade systems, which can be viewed as the limit, of a finite collection of ODE - $2 \times 2$ hyperbolic systems, as the number of individual PDE system…

Optimization and Control · Mathematics 2026-05-12 Jukka-Pekka Humaloja , Nikolaos Bekiaris-Liberis

A deep-learning-based closure model to address energy loss in low-dimensional surrogate models based on proper-orthogonal-decomposition (POD) modes is introduced. Using a transformer-encoder block with easy-attention mechanism, the model…

We study the gradient-based training of large-depth residual networks (ResNets) from standard random initializations. We show that infinite-depth ResNets behave as if they were infinitely wide, regardless of their actual width. More…

Machine Learning · Computer Science 2026-03-04 Lénaïc Chizat

Large-scale foundation models for scientific machine learning adapt to physical settings unseen during training, such as zero-shot transfer between turbulent scales. This phenomenon, in-context learning, challenges conventional…

Machine Learning · Computer Science 2026-04-14 Anthony Bao , Jeffrey Lai , William Gilpin

Model reduction for fluid flow simulation continues to be of great interest across a number of scientific and engineering fields. In a previous work [arXiv:2104.13962], we explored the use of Neural Ordinary Differential Equations (NODE) as…

Machine Learning · Computer Science 2021-07-07 Sourav Dutta , Peter Rivera-Casillas , Orie M. Cecil , Matthew W. Farthing , Emma Perracchione , Mario Putti

Is the standard weight decay in AdamW truly optimal? Although AdamW decouples weight decay from adaptive gradient scaling, a fundamental conflict remains: the Radial Tug-of-War. In deep learning, gradients tend to increase parameter norms…

Machine Learning · Computer Science 2026-02-06 Hao Chen , Jinghui Yuan , Hanmin Zhang

Masked diffusion models (MDMs), which leverage bidirectional attention and a denoising process, are narrowing the performance gap with autoregressive models (ARMs). However, their internal attention mechanisms remain under-explored. This…

Artificial Intelligence · Computer Science 2026-01-13 Pengcheng Huang , Tianming Liu , Zhenghao Liu , Yukun Yan , Shuo Wang , Tong Xiao , Zulong Chen , Maosong Sun
‹ Prev 1 2 3 10 Next ›