中文
相关论文

相关论文: Token embeddings violate the manifold hypothesis

200 篇论文

The internal representations learned by language models consistently exhibit striking geometric structure: calendar months organize into a circle, historical years form a smooth one-dimensional manifold, and cities' latitudes and longitudes…

机器学习 · 计算机科学 2026-02-27 Dhruva Karkada , Daniel J. Korchinski , Andres Nava , Matthieu Wyart , Yasaman Bahri

General-purpose language models are trained to produce varied natural language outputs, but for some tasks, like annotation or classification, we need more specific output formats. LLM systems increasingly support structured output, which…

计算与语言 · 计算机科学 2025-08-04 Sil Hamilton , David Mimno

Reliable uncertainty quantification (UQ) is essential for ensuring trustworthy downstream use of large language models, especially when they are deployed in decision-support and other knowledge-intensive applications. Model certainty can be…

计算与语言 · 计算机科学 2025-11-04 Autumn Toney-Wails , Ryan Wails

Statistical neurodynamics studies macroscopic behaviors of randomly connected neural networks. We consider a deep layered feedforward network where input signals are processed layer by layer. The manifold of input signals is embedded in a…

无序系统与神经网络 · 物理学 2018-08-23 Shun-ichi Amari , Ryo Karakida , Masafumi Oizumi

Tomographic locality is a principle commonly used in the program of finding axioms that pick out quantum theory within the landscape of possible theories. The principle asserts the sufficiency of local measurements for achieving a…

We discuss quantum network Bell nonlocality in a setting where the network structure is not fully known. More concretely, an honest user may trust their local network topology, but not the structure of the rest of the network, involving…

量子物理 · 物理学 2025-01-07 Sadra Boreiri , Tamas Krivachy , Pavel Sekatski , Antoine Girardin , Nicolas Brunner

Contextualized embeddings vary by context, even for the same token, and form a distribution in the embedding space. To analyze this distribution, we focus on the norm of the mean embedding and the variance of the embeddings. In this study,…

计算与语言 · 计算机科学 2024-12-18 Hiroaki Yamagiwa , Hidetoshi Shimodaira

This work studies the capabilities of a large language model (LLM) to understand paralinguistic aspects of speech without fine-tuning its weights. We utilize an end-to-end system with a speech encoder, which is trained to produce token…

音频与语音处理 · 电气工程与系统科学 2025-09-03 Wonjune Kang , Junteng Jia , Chunyang Wu , Wei Zhou , Egor Lakomkin , Yashesh Gaur , Leda Sari , Suyoun Kim , Ke Li , Jay Mahadeokar , Ozlem Kalinli

Large Language Models (LLMs) have achieved remarkable performance and received significant research interest. The enormous computational demands, however, hinder the local deployment on devices with limited resources. The current prevalent…

密码学与安全 · 计算机科学 2026-02-13 Yujie Gu , Richeng Jin , Xiaoyu Ji , Yier Jin , Wenyuan Xu

Manifold hypotheses are typically used for tasks such as dimensionality reduction, interpolation, or improving classification performance. In the less common problem of manifold estimation, the task is to characterize the geometric…

机器学习 · 计算机科学 2019-06-19 Bharathkumar Ramachandra , Benjamin Dutton , Ranga Raju Vatsavai

As Large Language Models (LLMs) increasingly appear in social science research (e.g., economics and marketing), it becomes crucial to assess how well these models replicate human behavior. In this work, using hypothesis testing, we present…

计算机与社会 · 计算机科学 2025-06-19 Harbin Hong , Sebastian Caldas , Liu Leqi

Fine-tuning LLM-based text embedders via contrastive learning maps inputs and outputs into a new representational space, discarding the LLM's output semantics. We propose LLM2Vec-Gen, a self-supervised alternative that instead produces…

We present models for embedding words in the context of surrounding words. Such models, which we refer to as token embeddings, represent the characteristics of a word that are specific to a given context, such as word sense, syntactic…

计算与语言 · 计算机科学 2017-06-13 Lifu Tu , Kevin Gimpel , Karen Livescu

In this paper, we discuss how pure mathematics and theoretical physics can be applied to the study of language models. Using set theory and analysis, we formulate mathematically rigorous definitions of language models, and introduce the…

计算与语言 · 计算机科学 2024-08-01 Wenzhe Yang

Many approaches in the field of machine learning and data analysis rely on the assumption that the observed data lies on lower-dimensional manifolds. This assumption has been verified empirically for many real data sets. To make use of this…

机器学习 · 计算机科学 2022-09-27 Erik Thordsen , Erich Schubert

Text anomaly detection is a critical task in natural language processing (NLP), with applications spanning fraud detection, misinformation identification, spam detection and content moderation, etc. Despite significant advances in large…

计算与语言 · 计算机科学 2025-07-17 Feng Xiao , Jicong Fan

\def\mon{S^3\stackrel{S^1}{\rightarrow}S^2} \def\inst{S^7\stackrel{S^3}{\rightarrow}S^4} \def\octo{S^{15}\stackrel{S^7}{\rightarrow}S^8} In semilocal theories, the vacuum manifold is fibered in a non-trivial way by the action of the gauge…

高能物理 - 理论 · 物理学 2016-09-06 Mark Hindmarsh , Richard Holman , Thomas W. Kephart , Tanmay Vachaspati

Let X be an irreducible smooth complex projective curve of genus g>2, and let x be a fixed point. A framed bundle is a pair (E,\phi), where E is a vector bundle over X, of rank r and degree d, and \phi:E_x\to C^r is a non-zero homomorphism.…

代数几何 · 数学 2015-05-13 Indranil Biswas , Tomas L. Gomez , Vicente Muñoz

Large Language Models (LLMs) possess latent multi-token prediction (MTP) abilities despite being trained only for next-token generation. We introduce ESP (Embedding-Space Probing), a simple and training-free MTP method that probes an LLM…

计算与语言 · 计算机科学 2026-05-29 Raghavv Goel , Mukul Gagrani , Mingu Lee , Chris Lott

This paper introduces a novel Bayesian learning model to explain the behavior of Large Language Models (LLMs), focusing on their core optimization metric of next token prediction. We develop a theoretical framework based on an ideal…

机器学习 · 计算机科学 2024-09-25 Siddhartha Dalal , Vishal Misra