Related papers: The Mean-Field Dynamics of Transformers
Self-attention and masked self-attention are at the heart of Transformers' outstanding success. Still, our mathematical understanding of attention, in particular of its Lipschitz properties - which are key when it comes to analyzing…
Conventionally, Earth system (e.g., weather and climate) forecasting relies on numerical simulation with complex physical models and are hence both expensive in computation and demanding on domain expertise. With the explosive growth of the…
Standard inference and training with transformer based architectures scale quadratically with input sequence length. This is prohibitively large for a variety of applications especially in web-page translation, query-answering etc.…
Mean-field treatment (MFT) is frequently applied to approximately predict the dynamics of quantum optics systems, to simplify the system Hamiltonian through neglecting certain modes that are driven strongly or couple weakly with other…
We compare the accuracy of two cluster extensions of Dynamical Mean-Field Theory in describing d-wave superconductors, using as a reference model a saddle-point t-J model which can be solved exactly in the thermodynamic limit and at the…
Multi-head attention empowers the recent success of transformers, the state-of-the-art models that have achieved remarkable success in sequence modeling and beyond. These attention mechanisms compute the pairwise dot products between the…
In this work, we introduce the Prototypical Transformer (ProtoFormer), a general and unified framework that approaches various motion tasks from a prototype perspective. ProtoFormer seamlessly integrates prototype learning with Transformer…
Transformer models are revolutionizing machine learning, but their inner workings remain mysterious. In this work, we present a new visualization technique designed to help researchers understand the self-attention mechanism in transformers…
Quantum cluster theories are a set of approaches for the theory of correlated and disordered lattice systems, which treat correlations within the cluster explicitly, and correlations at longer length scales either perturbatively or within a…
Various forms of sparse attention have been explored to mitigate the quadratic computational and memory cost of the attention mechanism in transformers. We study sparse transformers not through a lens of efficiency but rather in terms of…
The dot product attention mechanism, originally designed for natural language processing tasks, is a cornerstone of modern Transformers. It adeptly captures semantic relationships between word pairs in sentences by computing a similarity…
Transformer models have recently garnered significant attention in image restoration due to their ability to capture long-range pixel dependencies. However, long-range attention often results in computational overhead without practical…
The rise of transformers in vision tasks not only advances network backbone designs, but also starts a brand-new page to achieve end-to-end image recognition (e.g., object detection and panoptic segmentation). Originated from Natural…
A ubiquitous approach to obtain transferable machine learning-based models of potential energy surfaces for atomistic systems is to decompose the total energy into a sum of local atom-centred contributions. However, in many systems…
A striking clustering phenomenon in the antiferromagnetic Hamiltonian Mean-Field model has been previously reported. The numerically observed bicluster formation and stabilization is here fully explained by a non linear analysis of the…
Vision transformers have gained popularity recently, leading to the development of new vision backbones with improved features and consistent performance gains. However, these advancements are not solely attributable to novel feature…
Approximating a probability distribution using a set of particles is a fundamental problem in machine learning and statistics, with applications including clustering and quantization. Formally, we seek a weighted mixture of Dirac measures…
Vision Transformers are very popular nowadays due to their state-of-the-art performance in several computer vision tasks, such as image classification and action recognition. Although their performance has been greatly enhanced through…
Mean-field theories of the glass transition predict a phase transition to a dynamically arrested state, yet no such transition is observed in experiments or simulations of finite-dimensional systems. We resolve this long-standing…
We develop a novel approach to understand the phases of one-dimensional Bose-Hubbard models. We integrate the simplicity of the mean-field theory and the numerical power of the density matrix renormalization group method to build an…