中文
相关论文

相关论文: Causal Dimensionality of Transformer Representatio…

200 篇论文

We consider the Schramm-Loewner evolution (SLE$_\kappa$) for $\kappa \in (4,8)$, which is the regime that the curve is self-intersecting but not space-filling. We let ${\mathcal K}$ be the set of $\kappa \in (4,8)$ for which the adjacency…

概率论 · 数学 2026-05-06 Konstantinos Kavvadias , Jason Miller , Lukas Schoug

We investigate the asymptotic properties of Bayesian bivariate causal discovery for Gaussian Linear Structural Equation Models (SEMs) with heteroscedastic noise. We demonstrate that with purely observational data, the posterior distribution…

统计理论 · 数学 2026-03-30 Valentinian Lungu , Anish Dhir , Mark van der Wilk , Ioannis Kontoyiannis

A widely cited result by Dong et al. (2021) showed that Transformers built from self-attention alone, without skip connections or feed-forward layers, suffer from rapid rank collapse: all token representations converge to a single…

机器学习 · 计算机科学 2026-04-28 Giansalvo Cirrincione

Structural causal models (SCMs) allow us to investigate complex systems at multiple levels of resolution. The causal abstraction (CA) framework formalizes the mapping between high- and low-level SCMs. We address CA learning in a challenging…

机器学习 · 计算机科学 2025-06-03 Gabriele D'Acunto , Fabio Massimo Zennaro , Yorgos Felekis , Paolo Di Lorenzo

Class imbalance significantly degrades classification performance, yet its effects are rarely analyzed from a unified theoretical perspective. We propose a principled framework based on three fundamental scales: the imbalance coefficient…

机器学习 · 统计学 2026-01-08 Rose Yvette Bandolo Essomba , Ernest Fokoué

What structural inductive bias helps transformers reason over knowledge graphs? Through controlled ablations of a minimal transformer modification with four independently removable components (sparse adjacency masking, edge-type biases,…

Transformer architectures have been widely adopted for time series forecasting, yet whether the representational mechanisms that make them powerful in NLP actually engage on time series data remains unexplored. The persistent…

机器学习 · 计算机科学 2026-05-07 Alper Yıldırım

Motivated by the hypothesis that neural network representations encode abstract, interpretable features as linearly accessible, approximately orthogonal directions, sparse autoencoders (SAEs) have become a popular tool in interpretability.…

机器学习 · 计算机科学 2025-11-05 Valérie Costa , Thomas Fel , Ekdeep Singh Lubana , Bahareh Tolooshams , Demba Ba

Answering one of the main questions of [FHK14, Chapter 7], we show that there is a tight connection between the depth of a classifiable shallow theory $T$ and the Borel rank of the isomorphism relation $\cong^\kappa_T$ on its models of size…

逻辑 · 数学 2020-04-07 Francesco Mangraviti , Luca Motto Ros

Is there really much more to say about sparse autoencoders (SAEs)? Autoencoders in general, and SAEs in particular, represent deep architectures that are capable of modeling low-dimensional latent structure in data. Such structure could…

机器学习 · 计算机科学 2025-06-09 Yin Lu , Xuening Zhu , Tong He , David Wipf

Sparse autoencoders (SAEs) model the activations of a neural network as linear combinations of sparsely occurring directions of variation (latents). The ability of SAEs to reconstruct activations follows scaling laws w.r.t. the number of…

机器学习 · 计算机科学 2025-09-05 Eric J. Michaud , Liv Gorton , Tom McGrath

Sparse Autoencoders (SAEs) provide potentials for uncovering structured, human-interpretable representations in Large Language Models (LLMs), making them a crucial tool for transparent and controllable AI systems. We systematically analyze…

机器学习 · 计算机科学 2026-02-03 Jack Gallifant , Shan Chen , Kuleen Sasse , Hugo Aerts , Thomas Hartvigsen , Danielle S. Bitterman

Time series foundation models (TSFMs) are increasingly deployed in high-stakes domains, yet their internal representations remain opaque. We present the first application of sparse autoencoders (SAEs) to a TSFM, training TopK SAEs on…

机器学习 · 计算机科学 2026-03-12 Anurag Mishra

A new unitarization approach incorporated with chiral symmetry is established and applied to study the $\pi K$ elastic scatterings. We demonstrate that the $\kappa$ resonance exists, if the scattering length parameter in the I=1/2, J=0…

高能物理 - 唯象学 · 物理学 2008-11-26 H. Q. Zheng , Z. Y. Zhou , G. Y. Qin , Z. G. Xiao , J. J. Wang , N. Wu

We investigate the relationship between representation geometry and neural network performance. Analyzing 52 pretrained ImageNet models across 13 architecture families, we show that effective dimension -- an unsupervised geometric metric --…

机器学习 · 计算机科学 2026-03-04 Sumit Yadav

Sparse Autoencoders (SAEs) are a prominent tool in mechanistic interpretability (MI) for decomposing neural network activations into interpretable features. However, the aspiration to identify a canonical set of features is challenged by…

机器学习 · 计算机科学 2025-05-27 Xiangchen Song , Aashiq Muhamed , Yujia Zheng , Lingjing Kong , Zeyu Tang , Mona T. Diab , Virginia Smith , Kun Zhang

Selective conformal prediction can yield substantially tighter uncertainty sets when we can identify calibration examples that are exchangeable with the test example. In interventional settings, such as perturbation experiments in genomics,…

机器学习 · 计算机科学 2026-03-03 Amir Asiaee , Kavey Aryan , James P. Long

We present the first systematic study of Sparse Autoencoders (SAEs) on video representations. Standard SAEs decompose video into interpretable, monosemantic features but destroy temporal coherence: hard TopK selection produces unstable…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Atahan Dokme , Sriram Vishwanath

Latent diffusion models have emerged as the dominant framework for high-fidelity and efficient image generation, owing to their ability to learn diffusion processes in compact latent spaces. However, while previous research has focused…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Qifan Li , Xingyu Zhou , Jinhua Zhang , Weiyi You , Shuhang Gu

We report a systematic failure mode in predictive representation learning. Across 2695 neural network configurations trained to predict linear-Gaussian dynamics, the optimal encoder tracks the environment rather than the system it is meant…

机器学习 · 计算机科学 2026-05-07 Kejun Liu