Related papers: Variance Is Not Importance: Structural Analysis of…
Starting from three-dimensional nonlinear elasticity under the restriction of incompressibility, we derive reduced models to capture the behavior of strings in response to external forces. Our $\Gamma$-convergence analysis of the…
We introduce directional routing, a lightweight mechanism that gives each transformer attention head learned suppression directions controlled by a shared router, at 3.9% parameter cost. We train a 433M-parameter model alongside an…
We numerically investigate collective ordering and disordering effects for vortices in type-II superconductors interacting with square and triangular substrate arrays under a dc drive that is slowly rotated with respect to the fixed…
The elastic properties of the $B_1$-structured transition-metal nitrides and their carbide counterparts are studied using the {\it ab initio\} density functional perturbation theory. The linear response results of elastic constants are in…
Transformer-based pre-trained models with millions of parameters require large storage. Recent approaches tackle this shortcoming by training adapters, but these approaches still require a relatively large number of parameters. In this…
Transformers are a widespread and successful model architecture, particularly in Natural Language Processing (NLP) and Computer Vision (CV). The essential innovation of this architecture is the Attention Mechanism, which solves the problem…
The layered structure of tetragonal Ni(CN)2, consisting of square-planar Ni(CN)4 units linked in the a-b plane, with no true periodicity along the c-axis, is expected to show anisotropic compression on the application of pressure.…
We study the generation, nonlinear development and secondary instability of unsteady G\"ortler vortices and streaks in compressible boundary layers exposed to free-stream vortical disturbances and evolving over concave, flat and convex…
The sensitive dependence of electronic and thermoelectric properties of MoS$_2$ on the applied strain opens up a variety of applications in the emerging area of straintronics. Using first principles based density functional theory…
Cross-flow turbines harness kinetic energy in wind or moving water. Due to their unsteady fluid dynamics, it can be difficult to predict the interplay between aspects of rotor geometry and turbine performance. This study considers the…
Transformers serve as the foundational architecture for large language and video generation models, such as GPT, BERT, SORA and their successors. Empirical studies have demonstrated that real-world data and learning tasks exhibit…
Compressing large neural networks is an important step for their deployment in resource-constrained computational platforms. In this context, vector quantization is an appealing framework that expresses multiple parameters using a single…
We report the structural, vibrational and electrical transport properties up to 16 GPa of the 1T-TiTe2, a prominent layered 2D system, which is predicted to show a series of topologically trivial - nontrivial transitions under hydrostatic…
Transfer learning is a powerful technique for knowledge-sharing between different tasks. Recent work has found that the representations of models with certain invariances, such as to adversarial input perturbations, achieve higher…
Using rigorous constitutive linearization of second variation introduced in [6] we study weak stability of homogeneous deformation of the axially compressed circular cylindrical shell, regarded as a 3-dimensional hyperelastic body. We show…
Deep Convolutional Neural Networks (CNNs) have long been the architecture of choice for computer vision tasks. Recently, Transformer-based architectures like Vision Transformer (ViT) have matched or even surpassed ResNets for image…
This paper investigates deep neural network (DNN) compression from the perspective of compactly representing and storing trained parameters. We explore the previously overlooked opportunity of cross-layer architecture-agnostic…
Ribbons are a class of slender structures whose length, width, and thickness are widely separated from each other. This scale separation gives a ribbon unusual mechanical properties in athermal macroscopic settings, e.g. it can bend without…
Recently, deep learning-based image compression has made signifcant progresses, and has achieved better ratedistortion (R-D) performance than the latest traditional method, H.266/VVC, in both subjective metric and the more challenging…
In recent years, deep architectures have been used for transfer learning with state-of-the-art performance in many datasets. The properties of their features remain, however, largely unstudied under the transfer perspective. In this work,…