English
Related papers

Related papers: Polymorphism Is Rotation: Operational Mechanistic …

200 papers

Sparse autoencoders (SAEs) operationalise the linear representation hypothesis: they reconstruct model activations as sparse linear combinations of interpretable dictionary atoms, on the implicit assumption that activation space is well…

Machine Learning · Computer Science 2026-05-12 Eslam Zaher , Maciej Trzaskowski , Quan Nguyen , Fred Roosta

We unify the discrete Fourier transform (DFT), discrete cosine transform (DCT), Walsh-Hadamard, Haar wavelet, Karhunen-Lo\`eve transform, and several others along with their continuous counterparts (Fourier transform, Fourier series,…

Signal Processing · Electrical Eng. & Systems 2026-05-19 Mitchell A. Thornton

Accurate rotation estimation is at the heart of robot perception tasks such as visual odometry and object pose estimation. Deep neural networks have provided a new way to perform these tasks, and the choice of rotation representation is an…

Computer Vision and Pattern Recognition · Computer Science 2021-01-19 Valentin Peretroukhin , Matthew Giamou , David M. Rosen , W. Nicholas Greene , Nicholas Roy , Jonathan Kelly

We present a rotation-equivariant unsupervised learning framework for the sparse deconvolution of non-negative scalar fields defined on the unit sphere. Spherical signals with multiple peaks naturally arise in Diffusion MRI (dMRI), where…

Image and Video Processing · Electrical Eng. & Systems 2021-02-19 Axel Elaldi , Neel Dey , Heejong Kim , Guido Gerig

While transformers excel in many settings, their application in the field of automated planning is limited. Prior work like PlanGPT, a state-of-the-art decoder-only transformer, struggles with extrapolation from easy to hard planning…

Artificial Intelligence · Computer Science 2025-08-12 Markus Fritzsche , Elliot Gestrin , Jendrik Seipp

We exhibit two distinct renormalization scenarios in many-parameter families of piecewise isometries (PWI) of a rhombus. The rotational component, defined over the quadratic field $\mathbb{K}=\mathbb{Q}(\sqrt{5})$, is fixed. The…

Dynamical Systems · Mathematics 2015-08-25 John H Lowenstein , Franco Vivaldi

Context. Metamorphic Testing is recognised in IEEE/ISO software-testing standards and increasingly recommended for AI systems, but its progress is bottlenecked by metamorphic relation (MR) identification: existing approaches (structured…

Software Engineering · Computer Science 2026-05-19 Meng Li , Xiaohua Yang , Jie Liu , Shiyu Yan

We present a model for dihadron fragmentation functions, describing the fragmentation of a quark into two unpolarized hadrons. We tune the parameters of our model to the output of the PYTHIA event generator for two-hadron semi-inclusive…

High Energy Physics - Phenomenology · Physics 2008-11-26 Alessandro Bacchetta , Marco Radici

We derive a Lorentzian OPE inversion formula for the principal series of $sl(2,\mathbb{R})$. Unlike the standard Lorentzian inversion formula in higher dimensions, the formula described here only applies to fully crossing-symmetric…

High Energy Physics - Theory · Physics 2019-07-24 Dalimil Mazac

Supervised learning has become a cornerstone of modern machine learning, yet a comprehensive theory explaining its effectiveness remains elusive. Empirical phenomena, such as neural analogy-making and the linear representation hypothesis,…

Machine Learning · Computer Science 2025-02-26 Patrik Reizinger , Alice Bizeul , Attila Juhos , Julia E. Vogt , Randall Balestriero , Wieland Brendel , David Klindt

Deep learning methods have proven capable of recovering operators between high-dimensional spaces, such as solution maps of PDEs and similar objects in mathematical physics, from very few training samples. This phenomenon of data-efficiency…

Machine Learning · Computer Science 2025-12-11 T. Mitchell Roddenberry , Leo Tzou , Ivan Dokmanić , Maarten V. de Hoop , Richard G. Baraniuk

Deep learning employs multi-layer neural networks trained via the backpropagation algorithm. This approach has achieved success across many domains and relies on adaptive gradient methods such as the Adam optimizer. Sequence modeling…

Machine Learning · Computer Science 2025-07-16 Esmail Gumaan

Transformers pretrained via next token prediction learn to factor their world into parts, representing these factors in orthogonal subspaces of the residual stream. We formalize two representational hypotheses: (1) a representation in the…

The remarkable capability of Transformers to do reasoning and few-shot learning, without any fine-tuning, is widely conjectured to stem from their ability to implicitly simulate a multi-step algorithms -- such as gradient descent -- with…

Machine Learning · Computer Science 2024-10-14 Khashayar Gatmiry , Nikunj Saunshi , Sashank J. Reddi , Stefanie Jegelka , Sanjiv Kumar

Sparse autoencoders (SAEs) are a promising approach to interpreting the internal representations of transformer language models. However, SAEs are usually trained separately on each transformer layer, making it difficult to use them to…

Machine Learning · Computer Science 2025-02-25 Tim Lawson , Lucy Farnik , Conor Houghton , Laurence Aitchison

Invariance under symmetry is an important problem in machine learning. Our paper looks specifically at equivariant neural networks where transformations of inputs yield homomorphic transformations of outputs. Here, steerable CNNs have…

Machine Learning · Computer Science 2021-09-15 Daniel Franzen , Michael Wand

In deep learning, models typically reuse the same parameters for all inputs. Mixture of Experts (MoE) defies this and instead selects different parameters for each incoming example. The result is a sparsely-activated model -- with…

Machine Learning · Computer Science 2022-06-20 William Fedus , Barret Zoph , Noam Shazeer

Transformers exhibit proficiency in capturing long-range dependencies, whereas State Space Models (SSMs) facilitate linear-time sequence modeling. Notwithstanding their synergistic potential, the integration of these architectures presents…

Computation and Language · Computer Science 2025-06-19 Bingheng Wu , Jingze Shi , Yifan Wu , Nan Tang , Yuyu Luo

The congruential rule advanced by Graves for polarization basis transformation of the radar backscatter matrix is now often misinterpreted as an example of consimilarity transformation. However, consimilarity transformations imply a…

Optics · Physics 2013-09-02 David Bebbington , Laura Carrea

Diffusion Probabilistic Models (DPMs) have shown a powerful capacity of generating high-quality image samples. Recently, diffusion autoencoders (Diff-AE) have been proposed to explore DPMs for representation learning via autoencoding. Their…

Computer Vision and Pattern Recognition · Computer Science 2023-03-02 Zijian Zhang , Zhou Zhao , Zhijie Lin