English
Related papers

Related papers: Scale Determines Whether Language Models Organize …

200 papers

Evaluating the quality of learned representations without relying on a downstream task remains one of the challenges in representation learning. In this work, we present Geometric Component Analysis (GeomCA) algorithm that evaluates…

Machine Learning · Computer Science 2021-05-27 Petra Poklukar , Anastasia Varava , Danica Kragic

How do neural language models acquire a language's structure when trained for next-token prediction? We address this question by deriving theoretical scaling laws for neural network performance on synthetic datasets generated by the Random…

Machine Learning · Computer Science 2025-05-13 Francesco Cagnetta , Alessandro Favero , Antonio Sclocchi , Matthieu Wyart

Multilingual language models (LMs) organize representations for typologically and orthographically diverse languages into a shared parameter space, yet the nature of this internal organization remains elusive. In this work, we investigate…

Computation and Language · Computer Science 2026-04-22 Aastha A K Verma , Anwoy Chatterjee , Mehak Gupta , Tanmoy Chakraborty

We investigate the relationship between representation geometry and neural network performance. Analyzing 52 pretrained ImageNet models across 13 architecture families, we show that effective dimension -- an unsupervised geometric metric --…

Machine Learning · Computer Science 2026-03-04 Sumit Yadav

Scaling laws guide the development of large language models (LLMs) by offering estimates for the optimal balance of model size, tokens, and compute. More recently, loss-to-loss scaling laws that relate losses across pretraining datasets and…

Machine Learning · Computer Science 2026-05-21 Prasanna Mayilvahanan , Thaddäus Wiedemer , Sayak Mallick , Matthias Bethge , Wieland Brendel

Diffusion geometry is a manifold learning framework that uses random walks defined by Markov transition matrices to characterize the geometry of a dataset at multiple scales. We use diffusion geometry for neural representations,…

Machine Learning · Computer Science 2026-05-18 Atharva Khandait , Jan E. Gerken

Informally, the 'linear representation hypothesis' is the idea that high-level concepts are represented linearly as directions in some representation space. In this paper, we address two closely related questions: What does "linear…

Computation and Language · Computer Science 2026-05-18 Kiho Park , Yo Joong Choe , Victor Veitch

When a language model asserts that "the capital of Australia is Sydney," does it know this is wrong? We characterize the geometry of correctness representations across 9 models from 5 architecture families. The structure is simple: the…

Machine Learning · Computer Science 2026-02-10 Seonglae Cho , Zekun Wu , Kleyton Da Costa , Adriano Koshiyama

We present Entropic Mutual-Information Geometry Large-Language Model Alignment (ENIGMA), a novel approach to Large-Language Model (LLM) training that jointly improves reasoning, alignment and robustness by treating an organisation's…

Machine Learning · Computer Science 2025-10-17 Gareth Seneque , Lap-Hang Ho , Nafise Erfanian Saeedi , Jeffrey Molendijk , Ariel Kuperman , Tim Elson

Large scale neural models show impressive performance across a wide array of linguistic tasks. Despite this they remain, largely, black-boxes - inducing vector-representations of their input that prove difficult to interpret. This limits…

Computation and Language · Computer Science 2024-06-05 Henry Conklin , Kenny Smith

The body plan of the fruit fly is determined by the expression of just a handful of genes. We show that the spatial patterns of expression for several of these genes scale precisely with the size of the embryo. Concretely, discrete…

We show that the geometric relations between semantic features in large language models' hidden states closely mirror human psychological associations. We construct feature vectors corresponding to 360 words and project them on 32 semantic…

Computation and Language · Computer Science 2026-05-01 Austin C. Kozlowski , Andrei Boutyline

Neural models learn representations of high-dimensional data on low-dimensional manifolds. Multiple factors, including stochasticities in the training process, model architectures, and additional inductive biases, may induce different…

Machine Learning · Computer Science 2025-12-02 Hanlin Yu , Berfin Inal , Georgios Arvanitidis , Soren Hauberg , Francesco Locatello , Marco Fumero

We study how large language models (LLMs) ``think'' through their representation space. We propose a novel geometric framework that models an LLM's reasoning as flows -- embedding trajectories evolving where logic goes. We disentangle…

Artificial Intelligence · Computer Science 2026-03-05 Yufa Zhou , Yixiao Wang , Xunjian Yin , Shuyan Zhou , Anru R. Zhang

Understanding whether large language models (LLMs) capture structured meaning requires examining how they represent concept relationships. In this work, we study three models of increasing scale: Pythia-70M, GPT-2, and Llama 3.1 8B,…

Computation and Language · Computer Science 2026-04-01 Andor Diera , Ansgar Scherp

Language models exhibit strong robustness to paraphrasing, suggesting that semantic information may be encoded through stable internal representations, yet the structure and origin of such invariance remain unclear. We propose a local…

Machine Learning · Computer Science 2026-05-08 Agnibh Dasgupta , Abdullah Tanvir , Xin Zhong

Log-likelihood vectors define a common space for comparing language models as probability distributions, enabling unified comparisons across heterogeneous settings. We extend this framework to training checkpoints and intermediate layers,…

Computation and Language · Computer Science 2026-04-21 Ryo Kishino , Yusuke Takase , Momose Oyama , Hiroaki Yamagiwa , Hidetoshi Shimodaira

Word embeddings represent language vocabularies as clouds of $d$-dimensional points. We investigate how information is conveyed by the general shape of these clouds, instead of representing the semantic meaning of each token. Specifically,…

Computation and Language · Computer Science 2025-01-15 Ondřej Draganov , Steven Skiena

Starting from the hypothesis that knowledge in semantic space is organized along structured manifolds, we argue that this geometric structure renders the space explorable. By traversing it and using the resulting continuous representations…

Artificial Intelligence · Computer Science 2026-01-14 Mateusz Bystroński , Doheon Han , Nitesh V. Chawla , Tomasz Kajdanowicz

Understanding how Large Language Models (LLMs) perform complex reasoning and their failure mechanisms is a challenge in interpretability research. To provide a measurable geometric analysis perspective, we define the concept of the…

Artificial Intelligence · Computer Science 2025-09-29 Bo Li , Guanzhi Deng , Ronghao Chen , Junrong Yue , Shuo Zhang , Qinghua Zhao , Linqi Song , Lijie Wen