English
Related papers

Related papers: Token embeddings violate the manifold hypothesis

200 papers

The internal representations learned by language models consistently exhibit striking geometric structure: calendar months organize into a circle, historical years form a smooth one-dimensional manifold, and cities' latitudes and longitudes…

Machine Learning · Computer Science 2026-02-27 Dhruva Karkada , Daniel J. Korchinski , Andres Nava , Matthieu Wyart , Yasaman Bahri

General-purpose language models are trained to produce varied natural language outputs, but for some tasks, like annotation or classification, we need more specific output formats. LLM systems increasingly support structured output, which…

Computation and Language · Computer Science 2025-08-04 Sil Hamilton , David Mimno

Reliable uncertainty quantification (UQ) is essential for ensuring trustworthy downstream use of large language models, especially when they are deployed in decision-support and other knowledge-intensive applications. Model certainty can be…

Computation and Language · Computer Science 2025-11-04 Autumn Toney-Wails , Ryan Wails

Statistical neurodynamics studies macroscopic behaviors of randomly connected neural networks. We consider a deep layered feedforward network where input signals are processed layer by layer. The manifold of input signals is embedded in a…

Disordered Systems and Neural Networks · Physics 2018-08-23 Shun-ichi Amari , Ryo Karakida , Masafumi Oizumi

Tomographic locality is a principle commonly used in the program of finding axioms that pick out quantum theory within the landscape of possible theories. The principle asserts the sufficiency of local measurements for achieving a…

We discuss quantum network Bell nonlocality in a setting where the network structure is not fully known. More concretely, an honest user may trust their local network topology, but not the structure of the rest of the network, involving…

Quantum Physics · Physics 2025-01-07 Sadra Boreiri , Tamas Krivachy , Pavel Sekatski , Antoine Girardin , Nicolas Brunner

Contextualized embeddings vary by context, even for the same token, and form a distribution in the embedding space. To analyze this distribution, we focus on the norm of the mean embedding and the variance of the embeddings. In this study,…

Computation and Language · Computer Science 2024-12-18 Hiroaki Yamagiwa , Hidetoshi Shimodaira

This work studies the capabilities of a large language model (LLM) to understand paralinguistic aspects of speech without fine-tuning its weights. We utilize an end-to-end system with a speech encoder, which is trained to produce token…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-03 Wonjune Kang , Junteng Jia , Chunyang Wu , Wei Zhou , Egor Lakomkin , Yashesh Gaur , Leda Sari , Suyoun Kim , Ke Li , Jay Mahadeokar , Ozlem Kalinli

Large Language Models (LLMs) have achieved remarkable performance and received significant research interest. The enormous computational demands, however, hinder the local deployment on devices with limited resources. The current prevalent…

Cryptography and Security · Computer Science 2026-02-13 Yujie Gu , Richeng Jin , Xiaoyu Ji , Yier Jin , Wenyuan Xu

Manifold hypotheses are typically used for tasks such as dimensionality reduction, interpolation, or improving classification performance. In the less common problem of manifold estimation, the task is to characterize the geometric…

Machine Learning · Computer Science 2019-06-19 Bharathkumar Ramachandra , Benjamin Dutton , Ranga Raju Vatsavai

As Large Language Models (LLMs) increasingly appear in social science research (e.g., economics and marketing), it becomes crucial to assess how well these models replicate human behavior. In this work, using hypothesis testing, we present…

Computers and Society · Computer Science 2025-06-19 Harbin Hong , Sebastian Caldas , Liu Leqi

Fine-tuning LLM-based text embedders via contrastive learning maps inputs and outputs into a new representational space, discarding the LLM's output semantics. We propose LLM2Vec-Gen, a self-supervised alternative that instead produces…

Computation and Language · Computer Science 2026-04-03 Parishad BehnamGhader , Vaibhav Adlakha , Fabian David Schmidt , Nicolas Chapados , Marius Mosbach , Siva Reddy

We present models for embedding words in the context of surrounding words. Such models, which we refer to as token embeddings, represent the characteristics of a word that are specific to a given context, such as word sense, syntactic…

Computation and Language · Computer Science 2017-06-13 Lifu Tu , Kevin Gimpel , Karen Livescu

In this paper, we discuss how pure mathematics and theoretical physics can be applied to the study of language models. Using set theory and analysis, we formulate mathematically rigorous definitions of language models, and introduce the…

Computation and Language · Computer Science 2024-08-01 Wenzhe Yang

Many approaches in the field of machine learning and data analysis rely on the assumption that the observed data lies on lower-dimensional manifolds. This assumption has been verified empirically for many real data sets. To make use of this…

Machine Learning · Computer Science 2022-09-27 Erik Thordsen , Erich Schubert

Text anomaly detection is a critical task in natural language processing (NLP), with applications spanning fraud detection, misinformation identification, spam detection and content moderation, etc. Despite significant advances in large…

Computation and Language · Computer Science 2025-07-17 Feng Xiao , Jicong Fan

\def\mon{S^3\stackrel{S^1}{\rightarrow}S^2} \def\inst{S^7\stackrel{S^3}{\rightarrow}S^4} \def\octo{S^{15}\stackrel{S^7}{\rightarrow}S^8} In semilocal theories, the vacuum manifold is fibered in a non-trivial way by the action of the gauge…

High Energy Physics - Theory · Physics 2016-09-06 Mark Hindmarsh , Richard Holman , Thomas W. Kephart , Tanmay Vachaspati

Let X be an irreducible smooth complex projective curve of genus g>2, and let x be a fixed point. A framed bundle is a pair (E,\phi), where E is a vector bundle over X, of rank r and degree d, and \phi:E_x\to C^r is a non-zero homomorphism.…

Algebraic Geometry · Mathematics 2015-05-13 Indranil Biswas , Tomas L. Gomez , Vicente Muñoz

Large Language Models (LLMs) possess latent multi-token prediction (MTP) abilities despite being trained only for next-token generation. We introduce ESP (Embedding-Space Probing), a simple and training-free MTP method that probes an LLM…

Computation and Language · Computer Science 2026-05-29 Raghavv Goel , Mukul Gagrani , Mingu Lee , Chris Lott

This paper introduces a novel Bayesian learning model to explain the behavior of Large Language Models (LLMs), focusing on their core optimization metric of next token prediction. We develop a theoretical framework based on an ideal…

Machine Learning · Computer Science 2024-09-25 Siddhartha Dalal , Vishal Misra