Related papers: Localmax dynamics for attention in transformers an…
Attention models are typically learned by optimizing one of three standard loss functions that are variously called -- soft attention, hard attention, and latent variable marginal likelihood (LVML) attention. All three paradigms are…
In this paper we reconsider the constraints which are imposed by relativistic requirements to any model of dynamical reduction. We review the debate on the subject and we call attention on the fundamental contributions by Aharonov and…
We establish the local asymptotic normality (LAN) property for estimating a multidimensional parameter in the drift of a system of $N$ interacting particles observed over a fixed time horizon in a mean-field regime $N \rightarrow \infty$.…
We present a theoretical analysis of the performance of transformer with softmax attention in in-context learning with linear regression tasks. While the existing literature predominantly focuses on the convergence of transformers with…
In this paper, we study the simultaneous stability problem of a finite number of locally inter-connected linear subsystems under practical constraints, including asynchronous and aperiodic sampling, time-varying delays, and measurement…
The Transformer model architecture has become one of the most widely used in deep learning and the attention mechanism is at its core. The standard attention formulation uses a softmax operation applied to a scaled dot product between query…
Conditional copula models allow dependence structures to vary with observed covariates while preserving a separation between marginal behavior and association. We study the uniform asymptotic behavior of kernel-weighted local likelihood…
We propose a new class of random feature methods for linearizing softmax and Gaussian kernels called hybrid random features (HRFs) that automatically adapt the quality of kernel estimation to provide most accurate approximation in the…
Transformer-based models have demonstrated exceptional performance across diverse domains, becoming the state-of-the-art solution for addressing sequential machine learning problems. Even though we have a general understanding of the…
The maximum element of the vector output by the Softmax function approaches zero as the input vector size increases. Transformer-based language models rely on Softmax to compute attention scores, causing the attention distribution to…
We study the dynamics of gradient flow for training a multi-head softmax attention model for in-context learning of multi-task linear regression. We establish the global convergence of gradient flow under suitable choices of initialization.…
We study the discrete constrained saddle dynamics and their momentum variants for locating saddle points on manifolds. Under the assumption of exact unstable eigenvectors, we establish a local linear convergence of the discrete constrained…
This paper studies the projected saddle-point dynamics associated to a convex-concave function, which we term saddle function. The dynamics consists of gradient descent of the saddle function in variables corresponding to convexity and…
Localization of wave functions in disordered systems can be characterized by the Lyapunov exponent, which is zero in the extended phase and nonzero in the localized phase. Previous studies have shown that this exponent is an analytic…
A mode-coupling theory for the slow single-particle dynamics in fluids adsorbed in disordered porous media is derived, which complements previous work on the collective dynamics [V. Krakoviack, Phys. Rev. E 75, 031503 (2007)]. Its…
While attention has been empirically shown to improve model performance, it lacks a rigorous mathematical justification. This short paper establishes a novel connection between attention mechanisms and multinomial regression. Specifically,…
Transformer-based models have emerged as one of the most widely used architectures for natural language processing, natural language generation, and image generation. The size of the state-of-the-art models has increased steadily reaching…
Weakly supervised temporal action localization is a challenging vision task due to the absence of ground-truth temporal locations of actions in the training videos. With only video-level supervision during training, most existing methods…
We develop a feedback control framework for stabilizing the McKean-Vlasov PDE on the torus. Our goal is to steer the dynamics toward a prescribed stationary distribution or accelerate convergence to it using a time-dependent control…
This is a survey on the local structure about a fixed point of discrete finite-dimensional holomorphic dynamical systems, discussing in particular the existence of local topological conjugacies to normal forms, and the structure of local…