English
Related papers

Related papers: Variance Is Not Importance: Structural Analysis of…

200 papers

Efficient training and inference algorithms, such as low-rank adaption and model pruning, have shown impressive performance for learning Transformer-based large foundation models. However, due to the technical challenges of the non-convex…

Machine Learning · Computer Science 2024-06-26 Hongkang Li , Meng Wang , Shuai Zhang , Sijia Liu , Pin-Yu Chen

Large transformers have demonstrated remarkable success, making it necessary to compress these models to reduce inference costs while preserving their perfor-mance. Current compression algorithms prune transformers at fixed compression…

Machine Learning · Computer Science 2025-03-03 Yizhuo Ding , Ke Fan , Yikai Wang , Xinwei Sun , Yanwei Fu

Topologically interlocked structures are assemblies of interlocking blocks that hold together solely through contact. Such structures have been shown to exhibit high strength, energy dissipation, and crack arrest properties. Recent studies…

Numerical Analysis · Mathematics 2023-10-02 Ioannis Koureas , Mohit Pundir , Shai Feldfogel , David S. Kammer

In this paper, we study the possibility of designing non-trivial random CSP models by exploiting the intrinsic connection between structures and typical-case hardness. We show that constraint consistency, a notion that has been developed to…

Artificial Intelligence · Computer Science 2011-10-12 J. Culberson , Y. Gao

Token compression techniques have recently emerged as powerful tools for accelerating Vision Transformer (ViT) inference in computer vision. Due to the quadratic computational complexity with respect to the token sequence length, these…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Phat Nguyen , Ngai-Man Cheung

We show that deep neural networks trained across diverse tasks exhibit remarkably similar low-dimensional parametric subspaces. We provide the first large-scale empirical evidence that demonstrates that neural networks systematically…

Machine Learning · Computer Science 2025-12-09 Prakhar Kaushik , Shravan Chaudhari , Ankit Vaidya , Rama Chellappa , Alan Yuille

This paper presents a set of validation metrics for transmission network parameters that is applicable in both creation of synthetic power system test cases and validation of existing models. Using actual data from two real-world power…

Applications · Statistics 2017-06-12 Mir Hadi Athari , Zhifang Wang

Attention layers, as commonly used in transformers, form the backbone of modern deep learning, yet there is no mathematical description of their benefits and deficiencies as compared with other architectures. In this work we establish both…

Machine Learning · Computer Science 2023-11-17 Clayton Sanford , Daniel Hsu , Matus Telgarsky

Random Number Generators play a critical role in a number of important applications. In practice, statistical testing is employed to gather evidence that a generator indeed produces numbers that appear to be random. In this paper, we…

Computational Complexity · Computer Science 2010-03-25 Weiling Chang , Binxing Fang , Xiaochun Yun , Shupeng Wang , Xiangzhan Yu

Despite their central role in the success of foundational models and large-scale language modeling, the theoretical foundations governing the operation of Transformers remain only partially understood. Contemporary research has largely…

Machine Learning · Computer Science 2025-06-02 Sagar Ghosh , Kushal Bose , Swagatam Das

Tensor train (TT) decomposition is a powerful representation for high-order tensors, which has been successfully applied to various machine learning tasks in recent years. However, since the tensor product is not commutative, permutation of…

Numerical Analysis · Computer Science 2017-05-31 Qibin Zhao , Masashi Sugiyama , Andrzej Cichocki

Embedding layers in transformer-based NLP models typically account for the largest share of model parameters, scaling with vocabulary size but not yielding performance gains proportional to scale. We propose an alternative approach in which…

Computation and Language · Computer Science 2025-05-06 Henry Ndubuaku , Mouad Talhi

(abridged) Quasar absorption lines provide a precise test of the assumed constancy of the fundamental constants of physics. We have investigated potential changes in the fine-structure constant, alpha, and the proton-to-electron mass ratio,…

Cosmology and Nongalactic Astrophysics · Physics 2012-03-01 Julian A. King

We introduce a high-throughput platform that enables simultaneous, parallel testing of six bistable beams via programmable motion of a rotating disk. By prescribing harmonic angular dynamics, the platform explores the phase space of angular…

Soft Condensed Matter · Physics 2025-12-11 Eduardo Gutierrez-Prieto , Gilad Yakir , Pedro M. Reis

We show how transformers can be used to vastly simplify neural video compression. Previous methods have been relying on an increasing number of architectural biases and priors, including motion prediction and warping operations, resulting…

Computer Vision and Pattern Recognition · Computer Science 2022-10-13 Fabian Mentzer , George Toderici , David Minnen , Sung-Jin Hwang , Sergi Caelles , Mario Lucic , Eirikur Agustsson

The widespread adoption of transfer learning has revolutionized machine learning by enabling efficient adaptation of pre-trained models to new domains. However, the reliability of these adaptations remains poorly understood, particularly…

Machine Learning · Computer Science 2025-09-01 Prabhav Singh , Jessica Sorrell

The chemical flexibility of metal-organic frameworks (MOFs) offers an ideal platform to tune structure and composition for specific applications, from gas sensing to catalysis and from photoelectric conversion to energy storage. This…

Materials Science · Physics 2024-02-13 Joshua Edzards , Holger-Dietrich Saßnick , Julia Santana Andreo , Caterina Cocchi

Recently, state-of-the-art approaches for pruning large pre-trained models (LPMs) have demonstrated that the training-free removal of non-critical residual blocks in Transformers is viable for reducing model size, achieving results that…

Machine Learning · Computer Science 2025-01-20 J. Pablo Muñoz , Jinjie Yuan , Nilesh Jain

Compression aims to reduce the size of an input, while maintaining its relevant properties. For multi-parameter persistent homology, compression is a necessary step in any computational pipeline, since standard constructions lead to large…

Algebraic Topology · Mathematics 2022-08-17 Ulderico Fugacci , Michael Kerber , Alexander Rolle

In this paper we consider nonlinearly elastic, frame-indifferent, and singularly perturbed two-well models for materials undergoing solid-solid phase transitions in any space dimensions, and we perform a simultaneous passage to…

Analysis of PDEs · Mathematics 2020-05-11 Elisa Davoli , Manuel Friedrich