中文
相关论文

相关论文: Geometric Scaling of Bayesian Inference in LLMs

200 篇论文

Large language models (LLMs) generalize smoothly across continuous semantic spaces, yet strict logical reasoning demands the formation of discrete decision boundaries. Prevailing theories relying on linear isometric projections fail to…

机器学习 · 计算机科学 2026-03-26 Long Zhang , Dai-jun Lin , Wei-neng Chen

Large Language Models (LLMs) drive current AI breakthroughs despite very little being known about their internal representations. In this work, we propose to shed the light on LLMs inner mechanisms through the lens of geometry. In…

人工智能 · 计算机科学 2024-07-12 Randall Balestriero , Romain Cosentino , Sarath Shekkizhar

The geometric structure of latent representations in large language models (LLMs) is an active area of research, driven in part by its implications for model transparency and AI safety. Existing literature has focused mainly on general…

机器学习 · 计算机科学 2026-04-14 Benjamin J. Choi , Melanie Weber

While parameter-efficient fine-tuning methods like low-rank adaptation (LoRA) are standard for large language models, principled estimation of epistemic uncertainty remains challenging. Recent results in the LoRA regime suggest that…

机器学习 · 计算机科学 2026-05-29 Daniel Dold , Emanuel Sommer , Julius Kobialka , Oliver Dürr , David Rügamer

Under real-analytic assumptions on decoder-only Transformers, recent work shows that the map from discrete prompts to last-token hidden states is generically injective on finite prompt sets. We refine this picture: for each layer $\ell$ we…

机器学习 · 计算机科学 2025-11-20 Mikael von Strauss

Despite their empirical success, pushing Transformer architectures to extreme depth often leads to a paradoxical failure: representations become increasingly redundant, lose rank, and ultimately collapse. Existing explanations largely…

机器学习 · 计算机科学 2026-01-16 Haoran Su , Chenyu You

High-dimensional datasets are well-approximated by low-dimensional structures. Over the past decade, this empirical observation motivated the investigation of detection, measurement, and modeling techniques to exploit these low-dimensional…

统计理论 · 数学 2015-12-15 Mauro Maggioni , Stanislav Minsker , Nate Strawn

Next-token predictors often appear to develop internal representations of the latent world and its rules. The probabilistic nature of these models suggests a deep connection between the structure of the world and the geometry of probability…

机器学习 · 计算机科学 2026-03-18 Sasha Brenner , Thomas R. Knösche , Nico Scherf

Large Language Models (LLMs) frequently prioritize conflicting in-context information over pre-existing parametric memory, a phenomenon often termed sycophancy or compliance. However, the mechanistic realization of this behavior remains…

机器学习 · 计算机科学 2026-02-09 Long Zhang , Fangwei Lin

We investigate Bayes posterior distributions in high-dimensional generalized linear models (GLMs) under the proportional asymptotics regime, where the number of features and samples diverge at a comparable rate. Specifically, we…

统计理论 · 数学 2026-01-05 Manuel Sáenz , Pragya Sur

Analysis of word embedding properties to inform their use in downstream NLP tasks has largely been studied by assessing nearest neighbors. However, geometric properties of the continuous feature space contribute directly to the use of…

Large Language Models (LLMs) demonstrate strong few-shot generalization through in-context learning, yet their reasoning in dynamic and stochastic environments remains opaque. Prior studies mainly focus on static tasks and overlook the…

人工智能 · 计算机科学 2025-12-23 Jensen Zhang , Jing Yang , Keze Wang

Vision-language models encode continuous geometry that their text pathway fails to express: a 6,000-parameter linear probe extracts hand joint angles at 6.1 degrees MAE from frozen features, while the best text output achieves only 20.0…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Yakov Pyotr Shkolnikov

Preserving geometric structure is important in learning. We propose a unified class of geometry-aware architectures that interleave geometric updates between layers, where both projection layers and intrinsic exponential map updates arise…

机器学习 · 计算机科学 2026-02-04 Karthik Elamvazhuthi , Shiba Biswal , Kian Rosenblum , Arushi Katyal , Tianli Qu , Grady Ma , Rishi Sonthalia

Geometric properties of Transformer weights, particularly the unembedding matrix, have been widely useful in language model interpretability research. Yet, their utility for estimating downstream performance remains unclear. In this work,…

计算与语言 · 计算机科学 2026-02-25 Atharva Kulkarni , Jacob Mitchell Springer , Arjun Subramonian , Swabha Swayamdipta

The completion of a Euclidean distance matrix (EDM) from sparse and noisy observations is a fundamental challenge in signal processing, with applications in sensor network localization, acoustic room reconstruction, molecular conformation,…

信号处理 · 电气工程与系统科学 2026-02-02 Rohit Varma Chiluvuri , Santosh Nannuru

We propose practical deep Gaussian process models on Riemannian manifolds, similar in spirit to residual neural networks. With manifold-to-manifold hidden layers and an arbitrary last layer, they can model manifold- and scalar-valued…

机器学习 · 统计学 2025-03-03 Kacper Wyrwal , Andreas Krause , Viacheslav Borovitskiy

Deep generative models are tremendously successful in learning low-dimensional latent representations that well-describe the data. These representations, however, tend to much distort relationships between points, i.e. pairwise distances…

机器学习 · 计算机科学 2018-09-14 Tao Yang , Georgios Arvanitidis , Dongmei Fu , Xiaogang Li , Søren Hauberg

Large language models (LLMs) show remarkable capabilities across a variety of tasks. Despite the models only seeing text in training, several recent studies suggest that LLM representations implicitly capture aspects of the underlying…

计算与语言 · 计算机科学 2024-04-16 Yutaro Yamada , Yihan Bao , Andrew K. Lampinen , Jungo Kasai , Ilker Yildirim

Inference in deep Bayesian neural networks is only fully understood in the infinite-width limit, where the posterior flexibility afforded by increased depth washes out and the posterior predictive collapses to a shallow Gaussian process.…

机器学习 · 计算机科学 2022-05-03 Jacob A. Zavatone-Veth , Cengiz Pehlevan