English
Related papers

Related papers: Polymorphism Is Rotation: Operational Mechanistic …

200 papers

This paper studies how Transformer models with Rotary Position Embeddings (RoPE) develop emergent, wavelet-like properties that compensate for the positional encoding's theoretical limitations. Through an analysis spanning model scales,…

Machine Learning · Computer Science 2025-06-06 Valeria Ruscio , Umberto Nanni , Fabrizio Silvestri

Two germs of linear analytic differential systems $x^{k+1}Y^\prime=A(x)Y$ with a non resonant irregular singularity are analytically equivalent if and only if they have the same eigenvalues and equivalent collections of Stokes matrices. The…

Dynamical Systems · Mathematics 2016-05-23 Jean-François Gagnon , Christiane Rousseau

Privacy-preserving neural network (NN) inference can be achieved by utilizing homomorphic encryption (HE), which allows computations to be directly carried out over ciphertexts. Popular HE schemes are built over large polynomial rings. To…

Cryptography and Security · Computer Science 2025-08-15 Sajjad Akherati , Xinmiao Zhang

Random Reshuffling (RR) is an algorithm for minimizing finite-sum functions that utilizes iterative gradient descent steps in conjunction with data reshuffling. Often contrasted with its sibling Stochastic Gradient Descent (SGD), RR is…

Optimization and Control · Mathematics 2021-04-06 Konstantin Mishchenko , Ahmed Khaled , Peter Richtárik

We present a deep learning approach for analyzing two-dimensional scattering data of semiflexible polymers under external forces. In our framework, scattering functions are compressed into a three-dimensional latent space using a…

Soft Condensed Matter · Physics 2025-04-24 Lijie Ding , Chi-Huan Tung , Bobby G. Sumpter , Wei-Ren Chen , Changwoo Do

Algorithmic reasoning requires capabilities which are most naturally understood through recurrent models of computation, like the Turing machine. However, Transformer models, while lacking recurrence, are able to perform such reasoning…

Machine Learning · Computer Science 2023-05-03 Bingbin Liu , Jordan T. Ash , Surbhi Goel , Akshay Krishnamurthy , Cyril Zhang

Tomography can be used to reveal internal properties of a 3D object using any penetrating wave. Advanced tomographic imaging techniques, however, are vulnerable to both systematic and random errors associated with the experimental…

Numerical Analysis · Mathematics 2019-02-08 Anthony P. Austin , Zichao Wendy Di , Sven Leyffer , Stefan M. Wild

Stationary potential scattering admits a formulation in terms of the quantum dynamics generated by a non-Hermitian effective Hamiltonian. We use this formulation to give a proof of the reciprocity theorem in two and three dimensions that…

Quantum Physics · Physics 2025-12-04 Farhang Loran , Ali Mostafazadeh

Sparse autoencoders (SAEs) have proven useful in disentangling the opaque activations of neural networks, primarily large language models, into sets of interpretable features. However, adapting them to domains beyond language, such as…

Machine Learning · Computer Science 2025-11-13 Ege Erdogan , Ana Lucic

The Subtree Isomorphism problem asks whether a given tree is contained in another given tree. The problem is of fundamental importance and has been studied since the 1960s. For some variants, e.g., ordered trees, near-linear time algorithms…

Computational Complexity · Computer Science 2015-10-16 Amir Abboud , Arturs Backurs , Thomas Dueholm Hansen , Virginia Vassilevska Williams , Or Zamir

Trained transformer models have been found to implement interpretable procedures for tasks like arithmetic and associative recall, but little is understood about how the circuits that implement these procedures originate during training. To…

Machine Learning · Computer Science 2024-10-08 Ziqian Zhong , Jacob Andreas

In this paper, we investigate two stochastic perturbations of the metamorphosis equations of image analysis, in the geometrical context of the Euler-Poincar\'e theory. In the metamorphosis of images, the Lie group of diffeomorphisms deforms…

Computer Vision and Pattern Recognition · Computer Science 2017-11-21 Alexis Arnaudon , Darryl Holm , Stefan Sommer

We introduce a graded formulation of internal symbolic computation for transformers. The hidden space is endowed with a grading $V=\bigoplus_{g\in G}V_g$, and symbolic operations are realized as typed block maps (morphisms)…

Machine Learning · Computer Science 2025-11-25 Tony Shaska

In this paper, a work-optimal parallelization of Kostelec and Rockmore's well-known fast Fourier transform and its inverse on the three-dimensional rotation group SO(3) is designed, implemented, and tested. To this end, the sequential…

Distributed, Parallel, and Cluster Computing · Computer Science 2018-08-03 Denis-Michael Lux , Christian Wülker , Gregory S. Chirikjian

Simple image rotations significantly reduce the accuracy of deep neural networks. Moreover, training with all possible rotations increases the data set, which also increases the training duration. In this work, we address trainable rotation…

Computer Vision and Pattern Recognition · Computer Science 2021-01-19 Wolfgang Fuhl , Enkelejda Kasneci

Most existing learning-based methods for solving imaging inverse problems can be roughly divided into two classes: iterative algorithms, such as plug-and-play and diffusion methods leveraging pretrained denoisers, and unrolled architectures…

Image and Video Processing · Electrical Eng. & Systems 2026-03-31 Matthieu Terris , Samuel Hurault , Maxime Song , Julian Tachella

Transformers are increasingly adopted for modeling and forecasting time-series, yet their internal mechanisms remain poorly understood from a dynamical systems perspective. In contrast to classical autoregressive and state-space models,…

Machine Learning · Computer Science 2025-12-25 Gregory Duthé , Nikolaos Evangelou , Wei Liu , Ioannis G. Kevrekidis , Eleni Chatzi

Recent studies in interpretability have explored the inner workings of transformer models trained on tasks across various domains, often discovering that these networks naturally develop highly structured representations. When such…

We train a linear attention transformer on millions of masked-block matrix completion tasks: each prompt is masked low-rank matrix whose missing block may be (i) a scalar prediction target or (ii) an unseen kernel slice of Nystr\"om…

Machine Learning · Computer Science 2025-09-25 Patrick Lutz , Aditya Gangrade , Hadi Daneshmand , Venkatesh Saligrama

Many identities involving symmetric functions can be proved through bijective manipulations of tableaux. In this paper, we prove identities involving polysymmetric functions through bijections and sign-reversing involutions. In their paper…

Combinatorics · Mathematics 2025-10-15 Aditya Khanna
‹ Prev 1 4 5 6 7 8 10 Next ›