English
Related papers

Related papers: Uniform Scaling Limits in AdamW-Trained Transforme…

200 papers

Local M-smoothers are interesting and important signal and image processing techniques with many connections to other methods. In our paper we derive a family of partial differential equations (PDEs) that result in one, two, and three…

Image and Video Processing · Electrical Eng. & Systems 2020-07-28 Martin Welk , Joachim Weickert

Transformer models have achieved state-of-the-art results across a diverse range of domains. However, concern over the cost of training the attention mechanism to learn complex dependencies between distant inputs continues to grow. In…

This manuscript investigates the one-pass stochastic gradient descent (SGD) dynamics of a two-layer neural network trained on Gaussian data and labels generated by a similar, though not necessarily identical, target function. We rigorously…

Machine Learning · Statistics 2023-02-14 Luca Arnaboldi , Ludovic Stephan , Florent Krzakala , Bruno Loureiro

Anderson localization of particles -- the complete halt of wave transport through multiple scattering and phase coherence -- is a paradigmatic manifestation of quantum interference in disordered media. In three dimensions, the scaling…

We study the problem of learning Transformer-based sequence models with black-box access to their outputs. In this setting, a learner may adaptively query the oracle with any sequence of vectors and observe the output of the target…

Machine Learning · Computer Science 2026-05-05 Satwik Bhattamishra , Kulin Shah , Michael Hahn , Varun Kanade

Self-attention in Transformers relies on globally normalized softmax weights, causing all tokens to compete for influence at every layer. When composed across depth, this interaction pattern induces strong synchronization dynamics that…

Machine Learning · Computer Science 2026-05-26 Jingkun Liu , Yisong Yue , Max Welling , Yue Song

Entity tracking requires maintaining and updating latent states for entities and attributes over long sequences. Recent task-specific attention operators can compress deep Transformer stacks into a few layers by performing multi-hop state…

Machine Learning · Computer Science 2026-05-22 Hangyue Zhao , Paul Caillon , Erwan Fagnou , Alexandre Allauzen

This paper addresses the local stabilization problem for semilinear single-track vehicle models with distributed tire friction dynamics, represented as interconnections of ordinary differential equations (ODEs) and hyperbolic partial…

Systems and Control · Electrical Eng. & Systems 2026-02-10 Luigi Romano , Ole Morten Aamo , Miroslav Krstić , Jan Åslund , Erik Frisk

Building scalable models to learn from diverse, multimodal data remains an open challenge. For vision-language data, the dominant approaches are based on contrastive learning objectives that train a separate encoder for each modality. While…

Computer Vision and Pattern Recognition · Computer Science 2022-10-24 Xinyang Geng , Hao Liu , Lisa Lee , Dale Schuurmans , Sergey Levine , Pieter Abbeel

We present a theoretical analysis of the training process for a single-layer GAN fed by high-dimensional input data. The training dynamics of the proposed model at both microscopic and macroscopic scales can be exactly analyzed in the…

Machine Learning · Computer Science 2019-10-29 Chuang Wang , Hong Hu , Yue M. Lu

This work addresses two major issues of end-to-end learned image compression (LIC) based on deep neural networks: variable-rate learning where separate networks are required to generate compressed images with varying qualities, and the…

Image and Video Processing · Electrical Eng. & Systems 2024-04-10 Wei Jiang , Wei Wang , Songnan Li , Shan Liu

Sequence models face a fundamental tradeoff between memory capacity and computational efficiency. Transformers achieve expressive context modeling at quadratic cost, while linear attention and state-space models run in linear time by…

Machine Learning · Computer Science 2026-05-11 Yaxita Amin , Helen Zichen Li , Mengfan Zhang , Samet Ayhan

Most ordinary differential equation (ODE) models used to describe biological or physical systems must be solved approximately using numerical methods. Perniciously, even those solvers which seem sufficiently accurate for the forward…

Transformers have transformed modern machine learning, driving breakthroughs in computer vision, natural language processing, and robotics. At the core of their success lies the attention mechanism, which enables the modeling of global…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Hemanth Saratchandran , Simon Lucey

We use a stochastic series expansion quantum Monte Carlo method to study the phase diagram of the one-dimensional extended Hubbard model at half filling for small to intermediate values of the on-site (U) and nearest-neighbor (V)…

Strongly Correlated Electrons · Physics 2009-11-07 Pinaki Sengupta , Anders W. Sandvik , David K. Campbell

Solids undergoing a transition from order to disorder experience the proliferation of topological defects. The melting process generates transient quantum states. However, their dynamical nature with femtosecond lifetime hinders exploration…

Strongly Correlated Electrons · Physics 2023-12-15 Sanghun Lee , Eunseo Kim , Junho Bang , Jongho Park , Changyoung Kim , Dirk Wulferding , Doohee Cho

We consider the Kob-Andersen model, a cooperative lattice gas with kinetic constraints which has been widely analyzed in the physics literature in connection with the study of the liquid/glass transition. We consider the model in a finite…

Probability · Mathematics 2020-09-02 Fabio Martinelli , Assaf Shapira , Cristina Toninelli

Continuous-time neural processes are performant sequential decision-makers that are built by differential equations (DE). However, their expressive power when they are deployed on computers is bottlenecked by numerical DE solvers. This…

As power systems transition toward renewable-rich and inverter-dominated operations, accurate time-domain dynamic analysis becomes increasingly critical. Such analysis supports key operational tasks, including transient stability…

Artificial Intelligence · Computer Science 2026-04-17 Haoran Li , Lihao Mai , Chenhan Xiao , Erik Blasch , Yang Weng

Understanding the training dynamics of transformers is important to explain the impressive capabilities behind large language models. In this work, we study the dynamics of training a shallow transformer on a task of recognizing…

Machine Learning · Computer Science 2024-10-15 Hongru Yang , Bhavya Kailkhura , Zhangyang Wang , Yingbin Liang
‹ Prev 1 8 9 10 Next ›