Related papers: Uniform Scaling Limits in AdamW-Trained Transforme…
Spatiotemporal forecasting in physical systems, such as large-scale traffic networks, requires modeling a dual dynamic: continuous macroscopic rhythms and discrete, unpredictable microscopic shocks. While Neural Ordinary Differential…
This paper deals with the backstepping design of observer-based compensators for parabolic ODE-PDE-ODE systems. The latter consist of n coupled parabolic PDEs with distinct diffusion coefficients and spatially-varying coefficients, that are…
The training and generalization dynamics of the Transformer's core mechanism, namely the Attention mechanism, remain under-explored. Besides, existing analyses primarily focus on single-head attention. Inspired by the demonstrated benefits…
Increasing the size of a Transformer does not always lead to enhanced performance. This phenomenon cannot be explained by the empirical scaling laws. Furthermore, the model's enhanced performance is closely associated with its memorization…
Motivated by engineering applications of subsea installation by deepwater construction vessels in oil drilling, and of aid delivery by unmanned aerial vehicles in disaster relief, we develop output-feedback boundary control of…
We present a finite-size scaling analysis of the droplet condensation-evaporation transition of a lattice gas (in two and three dimensions) and a Lennard-Jones gas (in three dimensions) at fixed density. Parallel multicanonical simulations…
We develop a transformer-based sequence-to-sequence model that recovers scalar ordinary differential equations (ODEs) in symbolic form from irregularly sampled and noisy observations of a single solution trajectory. We demonstrate in…
Several recently introduced deep learning optimizers utilizing matrix-level preconditioning have shown promising speedups relative to the current dominant optimizer AdamW, particularly in relatively small-scale experiments. However, efforts…
In this work, we propose a Multi-Window Masked Autoencoder (MW-MAE) fitted with a novel Multi-Window Multi-Head Attention (MW-MHA) module that facilitates the modelling of local-global interactions in every decoder transformer block through…
Recent advances in large language models have demonstrated the effectiveness of length scaling during post-training, yet its potential in pre-training remains underexplored. We present the Parallel Hidden Decoding Transformer…
Tensor Attention extends traditional attention mechanisms by capturing high-order correlations across multiple modalities, addressing the limitations of classical matrix-based attention. Meanwhile, Rotary Position Embedding…
Unsupervised multivariate time series anomaly detection (UMTSAD) plays a critical role in various domains, including finance, networks, and sensor systems. In recent years, due to the outstanding performance of deep learning in general…
After almost half a century since the work of Anderson [Phys. Rev. {\bf 109}, 1492 (1958)], at present there is no well established theoretical framework for understanding the dynamics of interacting particles in the presence of disorder.…
This work is motivated by the engineering challenge of suppressing vibrations in turbine blades of aero engines, which often operate under extreme thermal conditions and high-Mach aerodynamic environments that give rise to complex vibration…
We study a constrained training regime for decoder-only Transformers in which the token interface is fixed, previously trained dense blocks are not reopened, and the active trainable parameter set is kept approximately constant as depth…
In this paper, we establish well-posedness of reflected McKean-Vlasov SDEs and their particle approximations in smooth non-convex domains. We prove convergence of the interacting particle system to the corresponding mean-field limit with…
We consider the Maxwell's equations with perfect electric conductor boundary conditions in three-dimensional unbounded domains which are the union of a bounded resonator and one or several semi-infinite waveguides. We are interested in the…
Deriving closed-form, analytical expressions for reduced-order models, and judiciously choosing the closures leading to them, has long been the strategy of choice for studying phase- and noise-induced transitions for agent-based models…
We propose a method for learning dynamical systems from high-dimensional empirical data that combines variational autoencoders and (spatio-)temporal attention within a framework designed to enforce certain scientifically-motivated…
Motivated by applications to mathematical biology, we study the averaging problem for slow-fast systems, {\em in the case in which the fast dynamics is a stochastic process with multiple invariant measures}. We consider both the case in…