Related papers: Variance Is Not Importance: Structural Analysis of…
Different transformer architectures implement identical linguistic computations via distinct connectivity patterns, yielding model imprinted ``computational fingerprints'' detectable through spectral analysis. Using graph signal processing…
During the study of resistive switching devices, researchers have found that the influence of the insertion layer cannot be ignored. Many reports have confirmed that the appropriate insertion layer can significantly improve the performance…
Pre-trained large-scale language models have increasingly demonstrated high accuracy on many natural language processing (NLP) tasks. However, the limited weight storage and computational speed on hardware platforms have impeded the…
High pressure x-ray diffraction up to 30 GPa and resonant emission x-ray spectroscopy and partial fluorescence yield x-ray absorption spectroscopy up to 52 GPa were used to study how the structural and electronic properties of UTe$_2$…
A set of curves or images of similar shape is an increasingly common functional data set collected in the sciences. Principal Component Analysis (PCA) is the most widely used technique to decompose variation in functional data. However, the…
We predict that vertical transport in heterostructures formed by twisted graphene layers can exhibit a unique bistability mechanism. Intrinsically bistable $I$-$V$ characteristics arise from resonant tunneling and interlayer charge…
Vision transformer has demonstrated promising performance on challenging computer vision tasks. However, directly training the vision transformers may yield unstable and sub-optimal results. Recent works propose to improve the performance…
The architecture of the brain is too complex to be intuitively surveyable without the use of compressed representations that project its variation into a compact, navigable space. The task is especially challenging with high-dimensional…
Coupling $N$ large $m$ minimal models and flowing to IR fixed points is a systematic way to build new classes of compact unitary 2d CFTs which are likely to be irrational, and potentially have a positive Virasoro twist gap above the…
The Transformer architecture has inarguably revolutionized deep learning, overtaking classical architectures like multi-layer perceptrons (MLPs) and convolutional neural networks (CNNs). At its core, the attention block differs in form and…
The transition from monolithic to multi-component neural architectures in advanced neural network controllers poses substantial challenges due to the high computational complexity of the latter. Conventional model compression techniques for…
This paper is devoted to the variational derivation of reduced models for elastic membranes with fracture under constraints on the determinant of the deformation gradient. We consider two physically relevant settings: the…
We conduct a systematic study of the approximation properties of Transformer for sequence modeling with long, sparse and complicated memory. We investigate the mechanisms through which different components of Transformer, such as the…
Transformer architectures are typically described in algorithmic and statistical terms, leaving their internal mechanics without a familiar structural language for researchers trained in physical theories. To bridge this gap, we develop a…
Transformers, composed of multiple self-attention layers, hold strong promises toward a generic learning primitive applicable to different data modalities, including the recent breakthroughs in computer vision achieving state-of-the-art…
We derive kinetic equations describing injection and transport of spin polarized carriers in organic semiconductors with hopping conductivity via an impurity level. The model predicts a strongly voltage dependent magnetoresistance, defined…
We present a numerical study of 2D random-bond Potts ferromagnets. The model is studied both below and above the critical value $Q_c=4$ which discriminates between second and first-order transitions in the pure system. Two geometries are…
It has been observed that representations learned by distinct neural networks conceal structural similarities when the models are trained under similar inductive biases. From a geometric perspective, identifying the classes of…
The turbulent/non-turbulent interface is analysed in a direct numerical simulation of a boundary layer in the range $Re_\theta=2800-6600$, with emphasis on the behaviour of the relatively large-scale fractal intermittent region. This…
Vision Transformers (ViT) have marked a paradigm shift in computer vision, outperforming state-of-the-art models across diverse tasks. However, their practical deployment is hampered by high computational and memory demands. This study…