中文
相关论文

相关论文: Finding Belief Geometries with Sparse Autoencoders

200 篇论文

Powerful sentence encoders trained for multiple languages are on the rise. These systems are capable of embedding a wide range of linguistic properties into vector representations. While explicit probing tasks can be used to verify the…

计算与语言 · 计算机科学 2021-09-22 Maarten De Raedt , Fréderic Godin , Pieter Buteneers , Chris Develder , Thomas Demeester

Transformers generate valid and diverse chemical structures, but little is known about the mechanisms that enable these models to capture the rules of molecular representation. We present a mechanistic analysis of autoregressive…

机器学习 · 计算机科学 2025-12-11 Kristof Varadi , Mark Marosi , Peter Antal

Embodied trajectories, such as the executable motion sequences of robotic manipulators, underwater vehicles, and mobile robots, are a fundamental output of embodied AI. Modern generative models often treat them as a dense, monolithic signal…

机器人学 · 计算机科学 2026-05-25 Yan Tang , Yuanbo Tang , Tingyu Cao , Shaolun Huang , Yang Li

Despite their impressive performance, generative image models trained on large-scale datasets frequently fail to produce images with seemingly simple concepts -- e.g., human hands or objects appearing in groups of four -- that are…

图形学 · 计算机科学 2025-06-25 Matyas Bohacek , Thomas Fel , Maneesh Agrawala , Ekdeep Singh Lubana

We show that singular value decomposition of the lm_head} weight matrix of a transformer-based large language model -- requiring only five lines of PyTorch and no model inference -- reveals interpretable semantic subspaces directly from the…

机器学习 · 计算机科学 2026-05-26 Hisashi Miyashita

Flow matching models generate high-fidelity molecular geometries but incur significant computational costs during inference, requiring hundreds of network evaluations. This inference overhead becomes the primary bottleneck when such models…

机器学习 · 计算机科学 2025-10-07 Johanna Sommer , John Rachwan , Nils Fleischmann , Stephan Günnemann , Bertrand Charpentier

Is a geometric model required to synthesize novel views from a single image? Being bound to local convolutions, CNNs need explicit 3D biases to model geometric transformations. In contrast, we demonstrate that a transformer-based model can…

计算机视觉与模式识别 · 计算机科学 2021-08-31 Robin Rombach , Patrick Esser , Björn Ommer

Deep learning has demonstrated significant potential in medical imaging; however, the opacity of "black-box" models hinders clinical trust, while segmentation tasks typically necessitate labourious, hard-to-obtain pixel-wise annotations. To…

图像与视频处理 · 电气工程与系统科学 2026-01-22 Soumick Chatterjee , Hadya Yassin , Florian Dubost , Andreas Nürnberger , Oliver Speck

This paper presents a unified geometric framework for the statistical analysis of a general ill-posed linear inverse model which includes as special cases noisy compressed sensing, sign vector recovery, trace regression, orthogonal matrix…

统计理论 · 数学 2020-07-27 T. Tony Cai , Tengyuan Liang , Alexander Rakhlin

Sparse Autoencoders (SAEs) have shown to find interpretable features in neural networks from polysemantic neurons caused by superposition. Previous work has shown SAEs are an effective tool to extract interpretable features from the early…

计算机视觉与模式识别 · 计算机科学 2025-04-17 Matthew Bozoukov

Pretrained encoders for mathematical texts have achieved significant improvements on various tasks such as formula classification and information retrieval. Yet they remain limited in representing and capturing student strategies for entire…

计算机与社会 · 计算机科学 2026-04-13 Siddhartha Pradhan , Ethan Prihar , Erin Ottmar

Latent representations learned by neural networks often exhibit semantic structure, where concept similarity is reflected by geometric proximity in embedding space. However, comparing such spaces across models remains difficult: changes in…

Sparse autoencoders (SAEs) have become a central tool for interpreting language models. However, two key SAE analyses that remain difficult to scale are (1) matching semantically similar features across multi-layers and (2) compressing…

机器学习 · 计算机科学 2026-05-28 Tue M. Cao , Nguyen Do , My T. Thai

Sparse Autoencoders (SAEs) are a powerful dictionary learning technique for decomposing neural network activations, translating the hidden state into human ideas with high semantic value despite no external intervention or guidance.…

机器学习 · 计算机科学 2025-12-17 Albert Miao , Chenliang Zhou , Jiawei Zhou , Cengiz Oztireli

Hidden Markov models (HMMs) are probabilistic functions of finite Markov chains, or, put in other words, state space models with finite state space. In this paper, we examine subspace estimation methods for HMMs whose output lies a finite…

统计理论 · 数学 2009-11-20 Sofia Andersson , Tobias Rydén

The manifold hypothesis (MH) is often used to explain how machine learning can overcome the curse of dimensionality. However, the MH is only applicable in regimes where the training data provides a sufficiently dense sample of the…

Understanding the internal representations of large language models is crucial for ensuring their reliability and safety, with sparse autoencoders (SAEs) emerging as a promising interpretability approach. However, current SAE training…

机器学习 · 计算机科学 2025-10-13 T. Ed Li , Junyu Ren

Analysis of word embedding properties to inform their use in downstream NLP tasks has largely been studied by assessing nearest neighbors. However, geometric properties of the continuous feature space contribute directly to the use of…

Utilizing patch-based transformers for unstructured geometric data such as polygon meshes presents significant challenges, primarily due to the absence of a canonical ordering and variations in input sizes. Prior approaches to handling 3D…

计算机视觉与模式识别 · 计算机科学 2024-11-04 Mohammad Farazi , Yalin Wang

We investigate the emergence of different effective geometries in stochastic Clifford circuits with sparse coupling. By changing the probability distribution for choosing two-site gates as a function of distance, we generate sparse…

量子物理 · 物理学 2022-06-13 Tomohiro Hashizume , Sridevi Kuriyattil , Andrew J. Daley , Gregory Bentsen