English
Related papers

Related papers: Geometric Scaling of Bayesian Inference in LLMs

200 papers

Large language models (LLMs) generalize smoothly across continuous semantic spaces, yet strict logical reasoning demands the formation of discrete decision boundaries. Prevailing theories relying on linear isometric projections fail to…

Machine Learning · Computer Science 2026-03-26 Long Zhang , Dai-jun Lin , Wei-neng Chen

Large Language Models (LLMs) drive current AI breakthroughs despite very little being known about their internal representations. In this work, we propose to shed the light on LLMs inner mechanisms through the lens of geometry. In…

Artificial Intelligence · Computer Science 2024-07-12 Randall Balestriero , Romain Cosentino , Sarath Shekkizhar

The geometric structure of latent representations in large language models (LLMs) is an active area of research, driven in part by its implications for model transparency and AI safety. Existing literature has focused mainly on general…

Machine Learning · Computer Science 2026-04-14 Benjamin J. Choi , Melanie Weber

While parameter-efficient fine-tuning methods like low-rank adaptation (LoRA) are standard for large language models, principled estimation of epistemic uncertainty remains challenging. Recent results in the LoRA regime suggest that…

Machine Learning · Computer Science 2026-05-29 Daniel Dold , Emanuel Sommer , Julius Kobialka , Oliver Dürr , David Rügamer

Under real-analytic assumptions on decoder-only Transformers, recent work shows that the map from discrete prompts to last-token hidden states is generically injective on finite prompt sets. We refine this picture: for each layer $\ell$ we…

Machine Learning · Computer Science 2025-11-20 Mikael von Strauss

Despite their empirical success, pushing Transformer architectures to extreme depth often leads to a paradoxical failure: representations become increasingly redundant, lose rank, and ultimately collapse. Existing explanations largely…

Machine Learning · Computer Science 2026-01-16 Haoran Su , Chenyu You

High-dimensional datasets are well-approximated by low-dimensional structures. Over the past decade, this empirical observation motivated the investigation of detection, measurement, and modeling techniques to exploit these low-dimensional…

Statistics Theory · Mathematics 2015-12-15 Mauro Maggioni , Stanislav Minsker , Nate Strawn

Next-token predictors often appear to develop internal representations of the latent world and its rules. The probabilistic nature of these models suggests a deep connection between the structure of the world and the geometry of probability…

Machine Learning · Computer Science 2026-03-18 Sasha Brenner , Thomas R. Knösche , Nico Scherf

Large Language Models (LLMs) frequently prioritize conflicting in-context information over pre-existing parametric memory, a phenomenon often termed sycophancy or compliance. However, the mechanistic realization of this behavior remains…

Machine Learning · Computer Science 2026-02-09 Long Zhang , Fangwei Lin

We investigate Bayes posterior distributions in high-dimensional generalized linear models (GLMs) under the proportional asymptotics regime, where the number of features and samples diverge at a comparable rate. Specifically, we…

Statistics Theory · Mathematics 2026-01-05 Manuel Sáenz , Pragya Sur

Analysis of word embedding properties to inform their use in downstream NLP tasks has largely been studied by assessing nearest neighbors. However, geometric properties of the continuous feature space contribute directly to the use of…

Computation and Language · Computer Science 2019-04-11 Brendan Whitaker , Denis Newman-Griffis , Aparajita Haldar , Hakan Ferhatosmanoglu , Eric Fosler-Lussier

Large Language Models (LLMs) demonstrate strong few-shot generalization through in-context learning, yet their reasoning in dynamic and stochastic environments remains opaque. Prior studies mainly focus on static tasks and overlook the…

Artificial Intelligence · Computer Science 2025-12-23 Jensen Zhang , Jing Yang , Keze Wang

Vision-language models encode continuous geometry that their text pathway fails to express: a 6,000-parameter linear probe extracts hand joint angles at 6.1 degrees MAE from frozen features, while the best text output achieves only 20.0…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Yakov Pyotr Shkolnikov

Preserving geometric structure is important in learning. We propose a unified class of geometry-aware architectures that interleave geometric updates between layers, where both projection layers and intrinsic exponential map updates arise…

Machine Learning · Computer Science 2026-02-04 Karthik Elamvazhuthi , Shiba Biswal , Kian Rosenblum , Arushi Katyal , Tianli Qu , Grady Ma , Rishi Sonthalia

Geometric properties of Transformer weights, particularly the unembedding matrix, have been widely useful in language model interpretability research. Yet, their utility for estimating downstream performance remains unclear. In this work,…

Computation and Language · Computer Science 2026-02-25 Atharva Kulkarni , Jacob Mitchell Springer , Arjun Subramonian , Swabha Swayamdipta

The completion of a Euclidean distance matrix (EDM) from sparse and noisy observations is a fundamental challenge in signal processing, with applications in sensor network localization, acoustic room reconstruction, molecular conformation,…

Signal Processing · Electrical Eng. & Systems 2026-02-02 Rohit Varma Chiluvuri , Santosh Nannuru

We propose practical deep Gaussian process models on Riemannian manifolds, similar in spirit to residual neural networks. With manifold-to-manifold hidden layers and an arbitrary last layer, they can model manifold- and scalar-valued…

Machine Learning · Statistics 2025-03-03 Kacper Wyrwal , Andreas Krause , Viacheslav Borovitskiy

Deep generative models are tremendously successful in learning low-dimensional latent representations that well-describe the data. These representations, however, tend to much distort relationships between points, i.e. pairwise distances…

Machine Learning · Computer Science 2018-09-14 Tao Yang , Georgios Arvanitidis , Dongmei Fu , Xiaogang Li , Søren Hauberg

Large language models (LLMs) show remarkable capabilities across a variety of tasks. Despite the models only seeing text in training, several recent studies suggest that LLM representations implicitly capture aspects of the underlying…

Computation and Language · Computer Science 2024-04-16 Yutaro Yamada , Yihan Bao , Andrew K. Lampinen , Jungo Kasai , Ilker Yildirim

Inference in deep Bayesian neural networks is only fully understood in the infinite-width limit, where the posterior flexibility afforded by increased depth washes out and the posterior predictive collapses to a shallow Gaussian process.…

Machine Learning · Computer Science 2022-05-03 Jacob A. Zavatone-Veth , Cengiz Pehlevan