Related papers: Variance Is Not Importance: Structural Analysis of…
Efficient training and inference algorithms, such as low-rank adaption and model pruning, have shown impressive performance for learning Transformer-based large foundation models. However, due to the technical challenges of the non-convex…
Large transformers have demonstrated remarkable success, making it necessary to compress these models to reduce inference costs while preserving their perfor-mance. Current compression algorithms prune transformers at fixed compression…
Topologically interlocked structures are assemblies of interlocking blocks that hold together solely through contact. Such structures have been shown to exhibit high strength, energy dissipation, and crack arrest properties. Recent studies…
In this paper, we study the possibility of designing non-trivial random CSP models by exploiting the intrinsic connection between structures and typical-case hardness. We show that constraint consistency, a notion that has been developed to…
Token compression techniques have recently emerged as powerful tools for accelerating Vision Transformer (ViT) inference in computer vision. Due to the quadratic computational complexity with respect to the token sequence length, these…
We show that deep neural networks trained across diverse tasks exhibit remarkably similar low-dimensional parametric subspaces. We provide the first large-scale empirical evidence that demonstrates that neural networks systematically…
This paper presents a set of validation metrics for transmission network parameters that is applicable in both creation of synthetic power system test cases and validation of existing models. Using actual data from two real-world power…
Attention layers, as commonly used in transformers, form the backbone of modern deep learning, yet there is no mathematical description of their benefits and deficiencies as compared with other architectures. In this work we establish both…
Random Number Generators play a critical role in a number of important applications. In practice, statistical testing is employed to gather evidence that a generator indeed produces numbers that appear to be random. In this paper, we…
Despite their central role in the success of foundational models and large-scale language modeling, the theoretical foundations governing the operation of Transformers remain only partially understood. Contemporary research has largely…
Tensor train (TT) decomposition is a powerful representation for high-order tensors, which has been successfully applied to various machine learning tasks in recent years. However, since the tensor product is not commutative, permutation of…
Embedding layers in transformer-based NLP models typically account for the largest share of model parameters, scaling with vocabulary size but not yielding performance gains proportional to scale. We propose an alternative approach in which…
(abridged) Quasar absorption lines provide a precise test of the assumed constancy of the fundamental constants of physics. We have investigated potential changes in the fine-structure constant, alpha, and the proton-to-electron mass ratio,…
We introduce a high-throughput platform that enables simultaneous, parallel testing of six bistable beams via programmable motion of a rotating disk. By prescribing harmonic angular dynamics, the platform explores the phase space of angular…
We show how transformers can be used to vastly simplify neural video compression. Previous methods have been relying on an increasing number of architectural biases and priors, including motion prediction and warping operations, resulting…
The widespread adoption of transfer learning has revolutionized machine learning by enabling efficient adaptation of pre-trained models to new domains. However, the reliability of these adaptations remains poorly understood, particularly…
The chemical flexibility of metal-organic frameworks (MOFs) offers an ideal platform to tune structure and composition for specific applications, from gas sensing to catalysis and from photoelectric conversion to energy storage. This…
Recently, state-of-the-art approaches for pruning large pre-trained models (LPMs) have demonstrated that the training-free removal of non-critical residual blocks in Transformers is viable for reducing model size, achieving results that…
Compression aims to reduce the size of an input, while maintaining its relevant properties. For multi-parameter persistent homology, compression is a necessary step in any computational pipeline, since standard constructions lead to large…
In this paper we consider nonlinearly elastic, frame-indifferent, and singularly perturbed two-well models for materials undergoing solid-solid phase transitions in any space dimensions, and we perform a simultaneous passage to…