Related papers: Uniform Scaling Limits in AdamW-Trained Transforme…
Attention based Transformer architecture has enabled significant advances in the field of natural language processing. In addition to new pre-training techniques, recent improvements crucially rely on working with a relatively larger…
Bond-disordered Anderson model in two dimensions on a square lattice is studied numerically near the band center by calculating density of states (DoS), multifractal properties of eigenstates and the localization length. DoS divergence at…
Adaptive optimizers with decoupled weight decay, such as AdamW, are the de facto standard for pre-training large transformer-based generative models. Yet the quadratic nature of the $\ell_2$ penalty embedded in weight decay drives all…
Learning models of dynamical systems with external inputs, which may be, for example, nonsmooth or piecewise, is crucial for studying complex phenomena and predicting future state evolution, which is essential for applications such as…
We study the long-term qualitative behavior of randomly perturbed dynamical systems. More specifically, we look at limit cycles of stochastic differential equations (SDE) with Markovian switching, in which the process switches at random…
We propose a data-driven framework for learning reduced-order moment dynamics from PDE-governed systems using Neural ODEs. In contrast to derivative-based methods like SINDy, which necessitate densely sampled data and are sensitive to…
The combination of numerical integration and deep learning, i.e., ODE-net, has been successfully employed in a variety of applications. In this work, we introduce inverse modified differential equations (IMDE) to contribute to the behaviour…
Very little is known about the training dynamics of adaptive gradient methods like Adam in deep learning. In this paper, we shed light on the behavior of these algorithms in the full-batch and sufficiently large batch settings.…
This paper studies how Transformer models with Rotary Position Embeddings (RoPE) develop emergent, wavelet-like properties that compensate for the positional encoding's theoretical limitations. Through an analysis spanning model scales,…
We study the hardmax limit of self-attention dynamics for token embeddings obtained in the zero-temperature ($\beta\to+\infty$) regime, and relate it to the finite-$\beta$ setting. In this limit, the update rule can be viewed as a…
Learning rate scheduling is essential in transformer training, where the final annealing plays a crucial role in getting the best performance. However, the mechanisms behind this cooldown phase, with its characteristic drop in loss, remain…
Transformer-based models have delivered impressive results on many tasks, particularly vision and language tasks. In many model training situations, conventional configurations are typically adopted. For example, we often set the base model…
Causal inference in continuous-time sequential decision problems is challenged by hidden confounders. We show that, in latent state-space models with time-varying interventions, observability of the latent dynamics from observed data is…
This paper presents a delay-adaptive boundary control scheme for a $2\times 2$ coupled linear hyperbolic PDE-ODE cascade system with an unknown and arbitrarily long input delay. To construct a nominal delay-compensated control law, assuming…
The remarkable success of the Adam in training neural networks has naturally led to the widespread use of its descent-ascent counterpart, Adam-DA, for solving zero-sum games. Despite its popularity in practice, a rigorous theoretical…
Active matter swarms -- collectives of self-propelled particles that could self-assemble, ferry microscopic cargo, or endow materials with dynamic properties -- remain hard to steer. In crowded systems, tracking or controlling individual…
End-to-end backpropagation couples all layers through a global error signal, enabling coordinated learning but requiring long-range credit assignment. Motivated by recent progress in blockwise self-supervised learning (BWSSL), we ask…
Representation learning for high-dimensional, complex physical systems aims to identify a low-dimensional intrinsic latent space, which is crucial for reduced-order modeling and modal analysis. To overcome the well-known Kolmogorov barrier,…
We present an analytical theory for the most subradiant modes in a finite one-dimensional emitter array coupled to either an ideal or a nonideal waveguide. Using an effective non-Hermitian Hamiltonian together with a Bragg-edge…
Laser-driven Bose-Einstein condensate of ultracold atoms loaded into a lossy high-finesse optical resonator exhibits critical behavior and, in the thermodynamic limit, a phase transition between stationary states of different symmetries.…