English
Related papers

Related papers: Variance Is Not Importance: Structural Analysis of…

200 papers

Starting from three-dimensional nonlinear elasticity under the restriction of incompressibility, we derive reduced models to capture the behavior of strings in response to external forces. Our $\Gamma$-convergence analysis of the…

Analysis of PDEs · Mathematics 2023-06-22 Dominik Engl , Carolin Kreisbeck

We introduce directional routing, a lightweight mechanism that gives each transformer attention head learned suppression directions controlled by a shared router, at 3.9% parameter cost. We train a 433M-parameter model alongside an…

Machine Learning · Computer Science 2026-03-17 Kevin Taylor

We numerically investigate collective ordering and disordering effects for vortices in type-II superconductors interacting with square and triangular substrate arrays under a dc drive that is slowly rotated with respect to the fixed…

Superconductivity · Physics 2015-05-27 C. Reichhardt , C. J. Olson Reichhardt

The elastic properties of the $B_1$-structured transition-metal nitrides and their carbide counterparts are studied using the {\it ab initio\} density functional perturbation theory. The linear response results of elastic constants are in…

Materials Science · Physics 2009-11-10 Zhigang Wu , Xiao-Jia Chen , Viktor V. Struzhkin , Ronald E. Cohen

Transformer-based pre-trained models with millions of parameters require large storage. Recent approaches tackle this shortcoming by training adapters, but these approaches still require a relatively large number of parameters. In this…

Computation and Language · Computer Science 2023-01-31 Chin-Lun Fu , Zih-Ching Chen , Yun-Ru Lee , Hung-yi Lee

Transformers are a widespread and successful model architecture, particularly in Natural Language Processing (NLP) and Computer Vision (CV). The essential innovation of this architecture is the Attention Mechanism, which solves the problem…

Machine Learning · Computer Science 2024-11-25 Bernhard Bermeitinger , Tomas Hrycej , Massimo Pavone , Julianus Kath , Siegfried Handschuh

The layered structure of tetragonal Ni(CN)2, consisting of square-planar Ni(CN)4 units linked in the a-b plane, with no true periodicity along the c-axis, is expected to show anisotropic compression on the application of pressure.…

We study the generation, nonlinear development and secondary instability of unsteady G\"ortler vortices and streaks in compressible boundary layers exposed to free-stream vortical disturbances and evolving over concave, flat and convex…

Fluid Dynamics · Physics 2024-12-11 Dongdong Xu , Pierre Ricco , Elena Marensi

The sensitive dependence of electronic and thermoelectric properties of MoS$_2$ on the applied strain opens up a variety of applications in the emerging area of straintronics. Using first principles based density functional theory…

Materials Science · Physics 2015-06-22 Swastibrata Bhattacharyya , Tribhuwan Pandey , Abhishek K. Singh

Cross-flow turbines harness kinetic energy in wind or moving water. Due to their unsteady fluid dynamics, it can be difficult to predict the interplay between aspects of rotor geometry and turbine performance. This study considers the…

Transformers serve as the foundational architecture for large language and video generation models, such as GPT, BERT, SORA and their successors. Empirical studies have demonstrated that real-world data and learning tasks exhibit…

Machine Learning · Computer Science 2026-05-19 Zhaiming Shen , Alex Havrilla , Rongjie Lai , Alexander Cloninger , Wenjing Liao

Compressing large neural networks is an important step for their deployment in resource-constrained computational platforms. In this context, vector quantization is an appealing framework that expresses multiple parameters using a single…

Computer Vision and Pattern Recognition · Computer Science 2021-04-13 Julieta Martinez , Jashan Shewakramani , Ting Wei Liu , Ioan Andrei Bârsan , Wenyuan Zeng , Raquel Urtasun

We report the structural, vibrational and electrical transport properties up to 16 GPa of the 1T-TiTe2, a prominent layered 2D system, which is predicted to show a series of topologically trivial - nontrivial transitions under hydrostatic…

Transfer learning is a powerful technique for knowledge-sharing between different tasks. Recent work has found that the representations of models with certain invariances, such as to adversarial input perturbations, achieve higher…

Machine Learning · Computer Science 2024-07-08 Till Speicher , Vedant Nanda , Krishna P. Gummadi

Using rigorous constitutive linearization of second variation introduced in [6] we study weak stability of homogeneous deformation of the axially compressed circular cylindrical shell, regarded as a 3-dimensional hyperelastic body. We show…

Analysis of PDEs · Mathematics 2013-01-28 Yury Grabovsky , Davit Harutyunyan

Deep Convolutional Neural Networks (CNNs) have long been the architecture of choice for computer vision tasks. Recently, Transformer-based architectures like Vision Transformer (ViT) have matched or even surpassed ResNets for image…

Computer Vision and Pattern Recognition · Computer Science 2021-10-11 Srinadh Bhojanapalli , Ayan Chakrabarti , Daniel Glasner , Daliang Li , Thomas Unterthiner , Andreas Veit

This paper investigates deep neural network (DNN) compression from the perspective of compactly representing and storing trained parameters. We explore the previously overlooked opportunity of cross-layer architecture-agnostic…

Computer Vision and Pattern Recognition · Computer Science 2021-11-22 Yuezhou Sun , Wenlong Zhao , Lijun Zhang , Xiao Liu , Hui Guan , Matei Zaharia

Ribbons are a class of slender structures whose length, width, and thickness are widely separated from each other. This scale separation gives a ribbon unusual mechanical properties in athermal macroscopic settings, e.g. it can bend without…

Statistical Mechanics · Physics 2021-12-28 Ee Hou Yong , Farisan Dary , Luca Giomi , L. Mahadevan

Recently, deep learning-based image compression has made signifcant progresses, and has achieved better ratedistortion (R-D) performance than the latest traditional method, H.266/VVC, in both subjective metric and the more challenging…

Image and Video Processing · Electrical Eng. & Systems 2022-06-23 Haisheng Fu , Feng Liang , Jie Liang , Binglin Li , Guohe Zhang , Jingning Han

In recent years, deep architectures have been used for transfer learning with state-of-the-art performance in many datasets. The properties of their features remain, however, largely unstudied under the transfer perspective. In this work,…

Computer Vision and Pattern Recognition · Computer Science 2016-11-17 Micael Carvalho , Matthieu Cord , Sandra Avila , Nicolas Thome , Eduardo Valle
‹ Prev 1 3 4 5 6 7 10 Next ›