English
Related papers

Related papers: Polymorphism Is Rotation: Operational Mechanistic …

200 papers

Symmetry in the parameter space of deep neural networks (DNNs) has proven beneficial for various deep learning applications. A well-known example is the permutation symmetry in Multi-Layer Perceptrons (MLPs), where permuting the rows of…

Machine Learning · Computer Science 2025-05-30 Binchi Zhang , Zaiyi Zheng , Zhengzhang Chen , Jundong Li

Large language models exhibit sophisticated capabilities, yet understanding how they work internally remains a central challenge. A fundamental obstacle is that training selects for behavior, not circuitry, so many weight configurations can…

Machine Learning · Computer Science 2026-02-27 Joshua S. Schiffman

Omnidirectional images and spherical representations of $3D$ shapes cannot be processed with conventional 2D convolutional neural networks (CNNs) as the unwrapping leads to large distortion. Using fast implementations of spherical and…

Computer Vision and Pattern Recognition · Computer Science 2020-12-09 Suhas Lohit , Shubhendu Trivedi

Random operators constitute fundamental building blocks of models of complex systems yet are far from fully understood. Here, we explain an asymmetry emerging upon repeating identical isotropic (uniformly random) operations. Specifically,…

Statistical Mechanics · Physics 2021-06-03 Malte Schröder , Marc Timme

We study time-reversal symmetry in dynamical systems with finite phase space, with applications to birational maps reduced over finite fields. For a polynomial automorphism with a single family of reversing symmetries, a universal (i.e.,…

Dynamical Systems · Mathematics 2015-05-13 John A. G. Roberts , Franco Vivaldi

One of basic difficulties of machine learning is handling unknown rotations of objects, for example in image recognition. A related problem is evaluation of similarity of shapes, for example of two chemical molecules, for which direct…

Machine Learning · Computer Science 2018-01-04 Jarek Duda

Self-supervised foundation models for digital pathology encode small patches from H\&E whole slide images into latent representations used for downstream tasks. However, the invariance of these representations to patch rotation remains…

Image and Video Processing · Electrical Eng. & Systems 2024-12-17 Matouš Elphick , Samra Turajlic , Guang Yang

A recursion operator is an integro-differential operator which maps a generalized symmetry of a nonlinear PDE to a new symmetry. Therefore, the existence of a recursion operator guarantees that the PDE has infinitely many higher-order…

Exactly Solvable and Integrable Systems · Physics 2013-01-08 D. E. Baldwin , W. Hereman

Partial differential equation (PDE) models and their associated variational energy formulations are often rotationally invariant by design. This ensures that a rotation of the input results in a corresponding rotation of the output, which…

Machine Learning · Computer Science 2022-03-21 Tobias Alt , Karl Schrader , Joachim Weickert , Pascal Peter , Matthias Augustin

Normal multi-scale transform [4] is a nonlinear multi-scale transform for representing geometric objects that has been recently investigated [1, 7, 10]. The restrictive role of the exact order of polynomial reproduction $P_e$ of the…

Numerical Analysis · Mathematics 2013-11-19 Stanislav Harizanov

We show that a constant number of self-attention layers can efficiently simulate, and be simulated by, a constant number of communication rounds of Massively Parallel Computation. As a consequence, we show that logarithmic depth is…

Machine Learning · Computer Science 2024-02-15 Clayton Sanford , Daniel Hsu , Matus Telgarsky

Reciprocal transformations mix the role of the dependent and independent variables to achieve simpler versions or even linearized versions of nonlinear PDEs. These transformations help in the identification of a plethora of PDEs available…

Mathematical Physics · Physics 2016-04-08 C. Sardon

Transformer architectures are typically described in algorithmic and statistical terms, leaving their internal mechanics without a familiar structural language for researchers trained in physical theories. To bridge this gap, we develop a…

Disordered Systems and Neural Networks · Physics 2026-03-18 Po-Hao Chang

In this paper, we study the problem of multivariate shuffled linear regression, where the correspondence between predictors and responses in a linear model is obfuscated by a latent permutation. Specifically, we investigate the model…

Machine Learning · Statistics 2026-03-24 Zhangsong Li

Transformers can generate predictions in two approaches: 1. auto-regressively by conditioning each sequence element on the previous ones, or 2. directly produce an output sequences in parallel. While research has mostly explored upon this…

Computer Vision and Pattern Recognition · Computer Science 2021-08-18 Andrea Alfieri , Yancong Lin , Jan C. van Gemert

Transformers are the dominant architecture in AI, yet why they work remains poorly understood. This paper offers a precise answer: a transformer is a Bayesian network. We establish this in five ways. First, we prove that every sigmoid…

Artificial Intelligence · Computer Science 2026-03-19 Gregory Coppola

We develop theory and software for rotation equivariant operators on scalar and vector fields, with diverse applications in simulation, optimization and machine learning. Rotation equivariance (covariance) means all fields in the system…

Machine Learning · Computer Science 2022-08-08 Paul Shen , Michael Herbst , Venkat Viswanathan

Data augmentation in feature space is effective to increase data diversity. Previous methods assume that different classes have the same covariance in their feature distributions. Thus, feature transform between different classes is…

Computer Vision and Pattern Recognition · Computer Science 2020-08-05 Yuke Zhu , Yan Bai , Yichen Wei

Embedding models trained separately on similar data often produce representations that encode stable information but are not directly interchangeable. This lack of interoperability raises challenges in several practical applications, such…

Machine Learning · Computer Science 2025-10-16 Lucas Maystre , Alvaro Ortega Gonzalez , Charles Park , Rares Dolga , Tudor Berariu , Yu Zhao , Kamil Ciosek

Sparse autoencoders (SAEs) are widely used to extract sparse, interpretable latents from transformer activations. We test whether commonly used SAE quality metrics and automatic explanation pipelines can distinguish trained transformers…

Machine Learning · Computer Science 2026-01-28 Thomas Heap , Tim Lawson , Lucy Farnik , Laurence Aitchison
‹ Prev 1 2 3 10 Next ›