English
Related papers

Related papers: Polymorphism Is Rotation: Operational Mechanistic …

200 papers

Transformers use the dense self-attention mechanism which gives a lot of flexibility for long-range connectivity. Over multiple layers of a deep transformer, the number of possible connectivity patterns increases exponentially. However,…

Machine Learning · Computer Science 2023-06-05 Md Shamim Hussain , Mohammed J. Zaki , Dharmashankar Subramanian

The success of machine learning algorithms is inherently related to the extraction of meaningful features, as they play a pivotal role in the performance of these algorithms. Central to this challenge is the quality of data representation.…

Image and Video Processing · Electrical Eng. & Systems 2025-07-10 Weronika Hryniewska-Guzik , Przemyslaw Biecek

Rotation invariance and translation invariance have great values in image recognition tasks. In this paper, we bring a new architecture in convolutional neural network (CNN) named cyclic convolutional layer to achieve rotation invariance in…

Computer Vision and Pattern Recognition · Computer Science 2017-06-20 Shiyuan Li

Self-supervised learning is a powerful paradigm for representation learning on unlabelled images. A wealth of effective new methods based on instance matching rely on data-augmentation to drive learning, and these have reached a rough…

Computer Vision and Pattern Recognition · Computer Science 2022-10-11 Linus Ericsson , Henry Gouk , Timothy M. Hospedales

Transformers are effective at inferring the latent task from context via two inference modes: recognizing a task seen during training, and adapting to a novel one. Recent interpretability studies have identified from middle-layer…

Machine Learning · Computer Science 2026-05-06 Hao Yan , Haolin Yang , Yiqiao Zhong

Sparse autoencoders are usually trained one layer at a time, even though transformer residual stream activations are strongly coupled across depth. This creates a practical problem for multi-layer interventions: different layerwise…

Machine Learning · Computer Science 2026-05-28 Prathyush Poduval , Calvin Yeung , Neel Desai , Mohsen Imani

Standard Transformers have a fixed computational depth, fundamentally limiting their ability to generalize to tasks requiring variable-depth reasoning, such as multi-hop graph traversal or nested logic. We propose a depth-recurrent…

Machine Learning · Computer Science 2026-03-24 Hung-Hsuan Chen

Transformers process tokens in parallel but are temporally shallow: at position $t$, each layer attends to key-value pairs computed based on the previous layer, yielding a depth capped by the number of layers. Recurrent models offer…

Machine Learning · Computer Science 2026-04-24 Costin-Andrei Oncescu , Depen Morwani , Samy Jelassi , Alexandru Meterez , Mujin Kwun , Sham Kakade

The Eigendecomposition of quadratic forms (symmetric matrices) guaranteed by the spectral theorem is a foundational result in applied mathematics. Motivated by a shared structure found in inferential problems of recent interest---namely…

Machine Learning · Computer Science 2018-02-26 Mikhail Belkin , Luis Rademacher , James Voss

Robot simulators are indispensable tools across many fields, and recent research has significantly improved their functionality by incorporating additional gradient information. However, existing differentiable robot simulators suffer from…

Robotics · Computer Science 2024-12-30 Xiaohan Ye , Xifeng Gao , Kui Wu , Zherong Pan , Taku Komura

In many contexts the modal properties of a structure change, either due to the impact of a changing environment, fatigue, or due to the presence of structural damage. For example during flight, an aircraft's modal properties are known to…

Machine Learning · Computer Science 2018-12-12 Prasad Cheema , Mehrisadat M. Alamdari , Gareth A. Vio

The Transformer architecture has revolutionized the field of sequence modeling and underpins the recent breakthroughs in large language models (LLMs). However, a comprehensive mathematical theory that explains its structure and operations…

Machine Learning · Computer Science 2026-04-14 Xue-Cheng Tai , Hao Liu , Lingfeng Li , Raymond H. Chan

A new isomorphism invariant of matroids is introduced, in the form of a quasisymmetric function. This invariant (1) defines a Hopf morphism from the Hopf algebra of matroids to the quasisymmetric functions, which is surjective if one uses…

Combinatorics · Mathematics 2020-06-01 Louis J. Billera , Ning Jia , Victor Reiner

In certain situations, neural networks are trained upon data that obey underlying symmetries. However, the predictions do not respect the symmetries exactly unless embedded in the network structure. In this work, we introduce architectures…

Machine Learning · Computer Science 2022-04-28 Anwesh Bhattacharya , Marios Mattheakis , Pavlos Protopapas

The object of this study is an integral operator $\mathcal{S}$ which averages functions in the Euclidean upper half-space $\mathbb{R}_{+}^{n}$ over the half-spheres centered on the topological boundary $\partial \mathbb{R}_{+}^{n}$. By…

Classical Analysis and ODEs · Mathematics 2009-10-09 Aleksei Beltukov

Recent advances in scanning tunneling and transmission electron microscopies (STM and STEM) have allowed routine generation of large volumes of imaging data containing information on the structure and functionality of materials. The…

Disordered Systems and Neural Networks · Physics 2021-06-24 Maxim Ziatdinov , Chun Yin Wong , Sergei V. Kalinin

Transformers are a type of neural network that have demonstrated remarkable performance across various domains, particularly in natural language processing tasks. Motivated by this success, research on the theoretical understanding of…

Machine Learning · Computer Science 2025-02-18 Naoki Takeshita , Masaaki Imaizumi

We develop a novel deep learning architecture for naturally complex-valued data, which is often subject to complex scaling ambiguity. We treat each sample as a field in the space of complex numbers. With the polar form of a complex-valued…

Computer Vision and Pattern Recognition · Computer Science 2019-06-25 Rudrasis Chakraborty , Jiayun Wang , Stella X. Yu

Integral transforms are invaluable mathematical tools to map functions into spaces where they are easier to characterize. We introduce the hyperdimensional transform as a new kind of integral transform. It converts square-integrable…

Machine Learning · Computer Science 2023-10-26 Pieter Dewulf , Michiel Stock , Bernard De Baets

We prove the (generalized) principal pivot transform is matrix monotone, in the sense of the L\"owner ordering, under minimal hypotheses. This improves on the recent results of J. E. Pascoe and R. Tully-Doyle, Monotonicity of the principal…

Functional Analysis · Mathematics 2023-02-13 Kenneth Beard , Aaron Welters