中文
相关论文

相关论文: Same Geometry, Opposite Noise: Transformer Magnitu…

200 篇论文

We propose a simple model for sample space reducing (SSR) stochastic process, where the dynamical variable denoting the size of the state space is continuous. In general, one can view the model as a multiplicative stochastic process, with a…

统计力学 · 物理学 2025-07-25 Rahul Chhimpa , Avinash Chand Yadav\

Linearization has emerged as a strategy for developing efficient language models (LMs). Starting from an existing Transformer-based LM, linearization replaces the attention component with computationally efficient subquadratic \textit{token…

计算与语言 · 计算机科学 2026-02-02 Patrick Haller , Jonas Golde , Alan Akbik

Though Large Language Models (LLMs) have shown remarkable abilities in mathematics reasoning, they are still struggling with performing numeric operations accurately, such as addition and multiplication. Numbers can be tokenized into tokens…

计算与语言 · 计算机科学 2024-09-30 Zhejian Zhou , Jiayu Wang , Dahua Lin , Kai Chen

Successive self-training on a language model's own outputs is widely characterized as a process of flattening: diversity drops, distributions narrow, and the text becomes "more like itself." We provide evidence that this characterization is…

计算与语言 · 计算机科学 2026-05-21 Ming Liu

A common feature of biological networks is the geometric property of self-similarity. Molecular regulatory networks through to circulatory systems, nervous systems, social systems and ecological trophic networks, show self-similar…

分子网络 · 定量生物学 2012-03-09 Simon DeDeo , David C. Krakauer

Adversarial perturbations are noise-like patterns that can subtly change the data, while failing an otherwise accurate classifier. In this paper, we propose to use such perturbations within a novel contrastive learning setup to build…

计算机视觉与模式识别 · 计算机科学 2020-04-17 Jue Wang , Anoop Cherian

This paper addresses the problem of detecting multidimensional subspace signals, which model range-spread targets, in noise of unknown covariance. It is assumed that a primary channel of measurements, possibly consisting of signal plus…

信号处理 · 电气工程与系统科学 2022-10-04 Danilo Orlando , Giuseppe Ricci , Louis L. Scharf

In this work, we investigate how Large Language Models (LLMs) adapt their internal representations when encountering inputs of increasing difficulty, quantified as the degree of out-of-distribution (OOD) shift. We reveal a consistent and…

The internal representations learned by language models consistently exhibit striking geometric structure: calendar months organize into a circle, historical years form a smooth one-dimensional manifold, and cities' latitudes and longitudes…

机器学习 · 计算机科学 2026-02-27 Dhruva Karkada , Daniel J. Korchinski , Andres Nava , Matthieu Wyart , Yasaman Bahri

Cosmic background neutrinos have a large velocity dispersion, which causes the evolution of long-wavelength density perturbations to depend on scale. This scale-dependent growth leads to the well-known suppression in the linear theory…

宇宙学与河外天体物理 · 物理学 2018-06-27 Chi-Ting Chiang , Wayne Hu , Yin Li , Marilena LoVerde

As pretrained large language models replace task-specific decoders in speech recognition, a critical question arises: do their text-derived priors make recognition fairer or more biased across demographic groups? We evaluate nine models…

计算与语言 · 计算机科学 2026-04-24 Srishti Ginjala , Eric Fosler-Lussier , Christopher W. Myers , Srinivasan Parthasarathy

Relationships in scientific data, such as the numerical and spatial distribution relations of features in univariate data, the scalar-value combinations' relations in multivariate data, and the association of volumes in time-varying and…

机器学习 · 计算机科学 2022-07-25 Xiangyang He , Yubo Tao , Shuoliu Yang , Haoran Dai , Hai Lin

The cognitive reality of irregular morphological patterns has been debated for decades: do speakers extend them to novel forms, or are they lexical artifacts? A neural network trained on distributional input offers a learnability test: if…

计算与语言 · 计算机科学 2026-05-26 Akhilesh Kakolu Ramarao , Kevin Tang , Dinah Baer-Henney

This paper investigates approximation-theoretic aspects of the in-context learning capability of the transformers in representing a family of noisy linear dynamical systems. Our first theoretical result establishes an upper bound on the…

机器学习 · 计算机科学 2025-10-22 Frank Cole , Yuxuan Zhao , Yulong Lu , Tianhao Zhang

Orthographic similarities across languages provide a strong signal for probabilistic decipherment, especially for closely related language pairs. The existing decipherment models, however, are not well-suited for exploiting these…

计算与语言 · 计算机科学 2015-08-11 Iftekhar Naim , Daniel Gildea

Sign language recognition suffers from catastrophic scaling failure: models achieving high accuracy on small vocabularies collapse at realistic sizes. Existing architectures treat signs as atomic visual patterns, learning flat…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Bryan Cheng , Austin Jin , Jasper Zhang

We numerically investigate hyperuniformity in two-dimensional frictionless jammed packings of bidisperse systems. Hyperuniformity is characterized by the suppression of density fluctuations at large length scales, and the structure factor…

软凝聚态物质 · 物理学 2025-07-18 Duc T. Dam , Takeshi Kawasaki , Atsushi Ikeda , Kunimasa Miyazaki

Since their introduction, Transformer architectures have dominated Natural Language Processing (NLP). However, recent research has highlighted an inherent anisotropy phenomenon in these models, presenting a significant challenge to their…

计算与语言 · 计算机科学 2026-04-13 Raphael Bernas , Fanny Jourdan , Antonin Poché , Céline Hudelot

Transformer-based table retrieval systems flatten structured tables into token sequences, making retrieval sensitive to the choice of serialization even when table semantics remain unchanged. We show that semantically equivalent…

计算与语言 · 计算机科学 2026-04-29 Kushal Raj Bhandari , Adarsh Singh , Jianxi Gao , Soham Dan , Vivek Gupta

Transformers trained in low precision can suffer forward-error amplification. We give a first-order, module-wise theory that predicts when and where errors grow. For self-attention we derive a per-layer bound that factorizes into three…

机器学习 · 计算机科学 2025-10-28 Jinwoo Baek