English
Related papers

Related papers: Clustering in Deep Stochastic Transformers

200 papers

Transformers are extremely successful machine learning models whose mathematical properties remain poorly understood. Here, we rigorously characterize the behavior of transformers with hardmax self-attention and normalization sublayers as…

Computation and Language · Computer Science 2026-05-14 Albert Alcalde , Giovanni Fantuzzi , Enrique Zuazua

The evolution of tokens through deep transformer models can be modeled as an interacting particle system that has been shown to exhibit an asymptotic clustering behavior akin to the synchronization phenomenon in Kuramoto models. In this…

Machine Learning · Computer Science 2026-05-12 Shi Chen , Zhengjiang Lin , Yury Polyanskiy , Philippe Rigollet

We propose a novel framework to perform classification via deep learning in the presence of noisy annotations. When trained on noisy labels, deep neural networks have been observed to first fit the training data with clean labels during an…

Machine Learning · Computer Science 2020-10-26 Sheng Liu , Jonathan Niles-Weed , Narges Razavian , Carlos Fernandez-Granda

Deep learning provides accurate collaborative filtering models to improve recommender system results. Deep matrix factorization and their related collaborative neural networks are the state-of-art in the field; nevertheless, both models…

Information Retrieval · Computer Science 2021-07-28 Jesús Bobadilla , Fernando Ortega , Abraham Gutiérrez , Ángel González-Prieto

Transformer-based diffusion models have demonstrated remarkable performance at generating high-quality samples. However, our theoretical understanding of the reasons for this success remains limited. For instance, existing models are…

Machine Learning · Computer Science 2026-04-14 Hongkang Li , Hancheng Min , Rene Vidal

We introduce a class of exactly solvable models which exhibit an ordering noise-induced phase transition driven by an entropic mechanism. In contrast with previous studies, order does not appear in this case as a result of an instability of…

Condensed Matter · Physics 2007-05-23 M. Ibanes , J. Garcia-Ojalvo , R. Toral , J. M. Sancho

Transformers with self-attention modules as their core components have become an integral architecture in modern large language and foundation models. In this paper, we study the evolution of tokens in deep encoder-only transformers at…

Analysis of PDEs · Mathematics 2026-05-12 Albert Alcalde , Leon Bungert , Konstantin Riedl , Tim Roith

We consider the dynamics of strongly localized systems subject to dephasing noise with arbitrary correlation time. Although noise inevitably induces delocalization, transport in the noise-induced delocalized phase is subdiffusive in a…

Disordered Systems and Neural Networks · Physics 2017-07-28 Sarang Gopalakrishnan , K. Ranjibul Islam , Michael Knap

Topology inference is a powerful tool to better understand the behaviours of network systems (NSs). Different from most of prior works, this paper is dedicated to inferring the directed topology of NSs from noisy observations, where the…

Systems and Control · Electrical Eng. & Systems 2025-09-03 Qing Jiao , Yushan Li , Jianping He

Topology optimization enables the design of highly efficient and complex structures, but conventional iterative methods, such as SIMP-based approaches, often suffer from high computational costs and sensitivity to initial conditions.…

Computational Engineering, Finance, and Science · Computer Science 2025-09-18 Aaron Lutheran , Srijan Das , Alireza Tabarraei

The Transformer architecture has revolutionized the field of sequence modeling and underpins the recent breakthroughs in large language models (LLMs). However, a comprehensive mathematical theory that explains its structure and operations…

Machine Learning · Computer Science 2026-04-14 Xue-Cheng Tai , Hao Liu , Lingfeng Li , Raymond H. Chan

We conjecture that the inherent difference in generalisation between adaptive and non-adaptive gradient methods in deep learning stems from the increased estimation noise in the flattest directions of the true loss surface. We demonstrate…

Machine Learning · Statistics 2022-03-17 Diego Granziol , Nicholas Baskerville

Although transformer-based models have shown exceptional empirical performance, the fundamental principles governing their training dynamics are inadequately characterized beyond configuration-specific studies. Inspired by empirical…

Machine Learning · Computer Science 2025-10-09 Zheng-An Chen , Tao Luo

We study how matrix-product-operator (MPO) truncation errors evolve when simulating two setups: (1) 1D Haar-random circuits under either depolarizing noise or amplitude-damping noise, and (2) 1D Lindbladian dynamics of a non-integrable…

Transformer architecture has shown impressive performance in multiple research domains and has become the backbone of many neural network models. However, there is limited understanding on how it works. In particular, with a simple…

Computation and Language · Computer Science 2023-10-31 Yuandong Tian , Yiping Wang , Beidi Chen , Simon Du

Factorized layers--operations parameterized by products of two or more matrices--occur in a variety of deep learning contexts, including compressed model training, certain types of knowledge distillation, and multi-head self-attention…

Machine Learning · Statistics 2022-10-07 Mikhail Khodak , Neil Tenenholtz , Lester Mackey , Nicolò Fusi

We show that the core components of the Transformer block -- attention, residual connections, and normalization -- arise naturally from a single geometric estimation problem. Modeling the latent state as a direction on the hypersphere, with…

Machine Learning · Computer Science 2026-05-13 Peter Racioppo

Self-propelled particles, like motile cells and artificial colloids, can spontaneously form macroscopic clusters. This phenomenon is called motility-induced phase separation (MIPS) and occurs even without attractive forces, provided that…

Soft Condensed Matter · Physics 2025-09-26 Felipe Hawthorne , Pablo de Castro , José A. Freire

Stochastic embedding transitions introduce a probabilistic mechanism for adjusting token representations dynamically during inference, mitigating the constraints imposed through static or deterministic embeddings. A transition framework was…

Computation and Language · Computer Science 2025-08-11 Stefan Whitaker , Colin Sisate , Marcel Windsor , Nikolai Fairweather , Tarquin Goldborough , Oskar Lindenfeld

Transformers process tokens in parallel but are temporally shallow: at position $t$, each layer attends to key-value pairs computed based on the previous layer, yielding a depth capped by the number of layers. Recurrent models offer…

Machine Learning · Computer Science 2026-04-24 Costin-Andrei Oncescu , Depen Morwani , Samy Jelassi , Alexandru Meterez , Mujin Kwun , Sham Kakade