Related papers: Variance Is Not Importance: Structural Analysis of…
Can a transformer learn which attention entries matter during training? In principle, yes: attention distributions are highly concentrated, and a small gate network can identify the important entries post-hoc with near-perfect accuracy. In…
Finetuning pretrained models occurs in a low-dimensional subspace of the full parameter space. Prior work has focused on characterizing this optimization subspace, but largely ignored the complementary question: why do certain directions…
Attention-based transformers have been remarkably successful at modeling generative processes across various domains and modalities. In this paper, we study the behavior of transformers on data drawn from \kth Markov processes, where the…
A lattice model of critical dense polymers is solved exactly for arbitrary system size on the torus. More generally, an infinite family of lattice loop models is studied on the torus and related to the corresponding Fortuin-Kasteleyn random…
Transformers have shown impressive capabilities across various tasks, but their performance on compositional problems remains a topic of debate. In this work, we investigate the mechanisms of how transformers behave on unseen compositional…
Modern compression systems use linear transformations in their encoding and decoding processes, with transforms providing compact signal representations. While multiple data-dependent transforms for image/video coding can adapt to diverse…
Many types of neural network layers rely on matrix properties such as invertibility or orthogonality. Retaining such properties during optimization with gradient-based stochastic optimizers is a challenging task, which is usually addressed…
We apply the model of stimulated neutrino transitions to neutrinos traveling through turbulence on a non-constant density profile. We describe a method to predict the location of large amplitude transitions and demonstrate the effectiveness…
Foundation models achieve state-of-the-art performance across different tasks, but their size and computational demands raise concerns about accessibility and sustainability. Existing efficiency methods often require additional retraining…
Compressing neural networks without retraining is vital for deployment at scale. We study calibration-free compression through the lens of projection geometry: structured pruning is an axis-aligned projection, whereas model folding performs…
With the popularity of the recent Transformer-based models represented by BERT, GPT-3 and ChatGPT, there has been state-of-the-art performance in a range of natural language processing tasks. However, the massive computations, huge memory…
Ninety eight one-dimensional channels defined using split gates fabricated on a GaAs/AlGaAs heterostructure are measured during one cooldown at 1.4 K. The devices are arranged in an array on a single chip, and individually addressed using a…
Interfaces have long been known to be the key to many mechanical and electric properties. To nickel base superalloys which have perfect creep and fatigue properties and have been widely used as materials of turbine blades, interfaces…
Scale invariance is a central organizing principle in physics, underlying phenomena that range from critical behaviour in statistical mechanics to transport and chaos in nonlinear dynamical systems. Here we present a unified and physically…
Ferromagnetic spintronics has been a main focus as it offers non-volatile memory and logic applications through current-induced spin-transfer torques. Enabling wider applications of such magnetic devices requires a lower switching current…
Two-body reduced density matrices (2RDMs) encode the essential two-electron physics of electronic states, but their quartic storage cost poses a major limitation in practical workflows. We investigate a simple protocol to compress both…
We train a linear attention transformer on millions of masked-block matrix completion tasks: each prompt is masked low-rank matrix whose missing block may be (i) a scalar prediction target or (ii) an unseen kernel slice of Nystr\"om…
We revisit the scaling properties of the resistivity and the current-voltage characteristics at and below the Berezinskii-Kosterlitz-Thouless transition, both in zero and nonzero magnetic field. The scaling properties are derived by…
The Transformer architecture has become the state-of-art model for natural language processing tasks and, more recently, also for computer vision tasks, thus defining the Vision Transformer (ViT) architecture. The key feature is the ability…
This paper presents a compression framework for Reservoir Computing that enables systematic design-space exploration of trade-offs among quantization levels, pruning rates, model accuracy, and hardware efficiency. The proposed approach…