Related papers: Variance Is Not Importance: Structural Analysis of…
Sparse neural networks are often hypothesized to be more interpretable than dense models, motivated by findings that weight sparsity can produce compact circuits in language models. However, it remains unclear whether structural sparsity…
A vision transformer (ViT) is the dominant model in the computer vision field. Despite numerous studies that mainly focus on dealing with inductive bias and complexity, there remains the problem of finding better transformer networks. For…
Variational auto-encoder frameworks have demonstrated success in reducing complex nonlinear dynamics in molecular simulation to a single non-linear embedding. In this work, we illustrate how this non-linear latent embedding can be used as a…
Motivation: Detailed experimental knowledge of the level structure of light weakly bound nuclei is necessary to guide the development of new theoretical approaches that combine nuclear structure with reaction dynamics. Purpose: The resonant…
In systems neuroscience, most models posit that brain regions communicate information under constraints of efficiency. Yet, evidence for efficient communication in structural brain networks characterized by hierarchical organization and…
Weight matrices in deep networks exhibit geometric continuity -- principal singular vectors of adjacent layers point in similar directions. While this property has been widely observed, its origin remains unexplained. Through experiments on…
Neural networks with self-attention (a.k.a. Transformers) like ViT and Swin have emerged as a better alternative to traditional convolutional neural networks (CNNs). However, our understanding of how the new architecture works is still…
We report three-dimensional particle mechanics static calculations that predict the microstructure evolution during die-compaction of elastic spherical particles up to relative densities close to one. We employ a nonlocal contact…
As a function of packing fraction at zero temperature and applied stress, an amorphous packing of spheres exhibits a jamming transition where the system is sensitive to boundary conditions even in the thermodynamic limit. Upon further…
Transients of a grid-synchronized voltage source converter (VSC) are closely related to over- currents and voltages occurred under large disturbances (e.g. a grid fault). Previous analysis in evaluating these transients usually neglect the…
Combined diverse two-dimensional (2D) materials for semiconductor interfaces are attractive for electrically controllable carrier confinement to enable excellent electrostatic control. We investigated the transport characteristic in…
The recently introduced structured input-output analysis is a powerful method for capturing nonlinear phenomena associated with incompressible flows, and this paper extends that method to the compressible regime. The proposed method relies…
In this work we present a study on the impact of various intrinsic deformations like ripples, twist, wrap on the electronic properties of ultra-short monolayer MoS2 channels. The effect of deformation (3-7o twist or wrap and 0.3-0.7…
In confined flows, such as river or tidal channels, arrays of turbines can convert both the kinetic and potential energy of the flow into renewable power. The power conversion and loading characteristics of an array in a confined flow is a…
Nonbenzenoid carbon frameworks expand low-dimensional material design via controlled asymmetry. Here, we show the experimentally realized 4-5-6-8 carbon nanoribbon establishes a topology-driven paradigm for multiproperty engineering, not…
It is now established that subcritical mechanisms play a crucial role in the transition to turbulence of non-rotating plane shear flows. The role of these mechanisms in rotating channel flow is examined here in the linear and nonlinear…
Transformers for language modeling usually rely on deterministic internal computation, with uncertainty expressed mainly at the output layer. We introduce variational neurons into Transformer feed-forward computation so that uncertainty…
Large-scale transformer training and deployment are increasingly constrained by the transfer of activations, gradients, and optimizer states across accelerators. Low-bit quantization offers a natural remedy, but transformer activations are…
Transformers have reshaped machine learning by utilizing attention mechanisms to capture complex patterns in large datasets, leading to significant improvements in performance. This success has contributed to the belief that "bigger means…
Heterogeneity is classified in five categories---topologic, geometric, kinematic, static, and constitutive---and the first four categories are investigated in a numerical DEM simulation of biaxial compression. The simulation experiments…