Related papers: Polymorphism Is Rotation: Operational Mechanistic …
The standard sparse-autoencoder (SAE) interpretability protocol labels each feature from its top-activating contexts and validates by single-feature steering. We propose the pairwise matrix protocol, co-varying steering coefficient with…
Solutions to the phonon Boltzmann transport equation under the relaxation-time approximation (RTA) are fundamentally limited in that they do not account for the off-diagonal elements of the scattering matrix, which encode intermode energy…
Panoramic semantic segmentation models are typically trained under a strict gravity-aligned assumption. However, real-world captures often deviate from this canonical orientation due to unconstrained camera motions, such as the rotational…
We study independent sets in strong powers of circulant graphs using a transfer matrix formulation. The compatibility constraints separate into intra-layer and inter-layer components, yielding a transfer operator that is equivariant under…
Paradoxically, a Variational Autoencoder (VAE) could be pushed in two opposite directions, utilizing powerful decoder model for generating realistic images but collapsing the learned representation, or increasing regularization coefficient…
In this paper, we contend that the objective of representation learning is to compress and transform the distribution of the data, say sets of tokens, towards a mixture of low-dimensional Gaussian distributions supported on incoherent…
This article demonstrates that convolutional operation can be converted to matrix multiplication, which has the same calculation way with fully connected layer. The article is helpful for the beginners of the neural network to understand…
We study the ability of Transformer models to learn sequences generated by Permuted Congruential Generators (PCGs), a widely used family of pseudo-random number generators (PRNGs). PCGs introduce substantial additional difficulty over…
The $S$ matrix rephasing invariance is one of the fundamental principles of quantum mechanics that originates in its probabilistic interpretation. For a given $S$ matrix which describes neutrino oscillation, one can define the two different…
Transformer-based models excel in various tasks but their generalization capabilities, especially in arithmetic reasoning, remain incompletely understood. Arithmetic tasks provide a controlled framework to explore these capabilities, yet…
Random Matrix Theory (RMT) is capable of making predictions for the spectral fluctuations of a physical system only after removing the influence of the level density by unfolding the spectra. When the level density is known, unfolding is…
Neural networks have emerged as effective tools for solving ill-posed inverse problems. In many scientific applications, however, observational training data are insufficient, and learned inverse operators must instead be trained on…
Collaborative perception in autonomous driving significantly enhances the perception capabilities of individual agents. Immutable heterogeneity, where agents have different and fixed perception networks, presents a major challenge due to…
Symmetries corresponding to local transformations of the fundamental fields that leave the action invariant give rise to (invertible) topological defects, which obey group-like fusion rules. One can construct more general (codimension-one)…
The perturbed GUE corners ensemble is the joint distribution of eigenvalues of all principal submatrices of a matrix $G+\mathrm{diag}(\mathbf{a})$, where $G$ is the random matrix from the Gaussian Unitary Ensemble (GUE), and…
Matrix functions extend scalar function concepts to linear operators, offering a unified framework with broad applications in mathematics, science, and engineering. Classical definitions--via power series, spectral calculus, or Jordan…
While transformers have proven enormously successful in a range of tasks, their fundamental properties as models of computation are not well understood. This paper contributes to the study of the expressive capacity of transformers,…
Single-operator learning involves training a deep neural network to learn a specific operator, whereas recent work in multi-operator learning uses an operator embedding structure to train a single neural network on data from multiple…
We generalize the supersymmetry method in Random Matrix Theory to arbitrary rotation invariant ensembles. Our exact approach further extends a previous contribution in which we constructed a supersymmetric representation for the class of…
In this paper we develop a new approach to the calibration of polarimetric radar data based on two key ideas. The first is the use of in-scene trihedral corner reflectors not only for radiometric and geometric calibration but also to…