English
Related papers

Related papers: Same Geometry, Opposite Noise: Transformer Magnitu…

200 papers

We propose a simple model for sample space reducing (SSR) stochastic process, where the dynamical variable denoting the size of the state space is continuous. In general, one can view the model as a multiplicative stochastic process, with a…

Statistical Mechanics · Physics 2025-07-25 Rahul Chhimpa , Avinash Chand Yadav\

Linearization has emerged as a strategy for developing efficient language models (LMs). Starting from an existing Transformer-based LM, linearization replaces the attention component with computationally efficient subquadratic \textit{token…

Computation and Language · Computer Science 2026-02-02 Patrick Haller , Jonas Golde , Alan Akbik

Though Large Language Models (LLMs) have shown remarkable abilities in mathematics reasoning, they are still struggling with performing numeric operations accurately, such as addition and multiplication. Numbers can be tokenized into tokens…

Computation and Language · Computer Science 2024-09-30 Zhejian Zhou , Jiayu Wang , Dahua Lin , Kai Chen

Successive self-training on a language model's own outputs is widely characterized as a process of flattening: diversity drops, distributions narrow, and the text becomes "more like itself." We provide evidence that this characterization is…

Computation and Language · Computer Science 2026-05-21 Ming Liu

A common feature of biological networks is the geometric property of self-similarity. Molecular regulatory networks through to circulatory systems, nervous systems, social systems and ecological trophic networks, show self-similar…

Molecular Networks · Quantitative Biology 2012-03-09 Simon DeDeo , David C. Krakauer

Adversarial perturbations are noise-like patterns that can subtly change the data, while failing an otherwise accurate classifier. In this paper, we propose to use such perturbations within a novel contrastive learning setup to build…

Computer Vision and Pattern Recognition · Computer Science 2020-04-17 Jue Wang , Anoop Cherian

This paper addresses the problem of detecting multidimensional subspace signals, which model range-spread targets, in noise of unknown covariance. It is assumed that a primary channel of measurements, possibly consisting of signal plus…

Signal Processing · Electrical Eng. & Systems 2022-10-04 Danilo Orlando , Giuseppe Ricci , Louis L. Scharf

In this work, we investigate how Large Language Models (LLMs) adapt their internal representations when encountering inputs of increasing difficulty, quantified as the degree of out-of-distribution (OOD) shift. We reveal a consistent and…

Computation and Language · Computer Science 2026-03-20 Mingyu Jin , Yutong Yin , Jingcheng Niu , Qingcheng Zeng , Wujiang Xu , Mengnan Du , Wei Cheng , Zhaoran Wang , Tianlong Chen , Dimitris N. Metaxas

The internal representations learned by language models consistently exhibit striking geometric structure: calendar months organize into a circle, historical years form a smooth one-dimensional manifold, and cities' latitudes and longitudes…

Machine Learning · Computer Science 2026-02-27 Dhruva Karkada , Daniel J. Korchinski , Andres Nava , Matthieu Wyart , Yasaman Bahri

Cosmic background neutrinos have a large velocity dispersion, which causes the evolution of long-wavelength density perturbations to depend on scale. This scale-dependent growth leads to the well-known suppression in the linear theory…

Cosmology and Nongalactic Astrophysics · Physics 2018-06-27 Chi-Ting Chiang , Wayne Hu , Yin Li , Marilena LoVerde

As pretrained large language models replace task-specific decoders in speech recognition, a critical question arises: do their text-derived priors make recognition fairer or more biased across demographic groups? We evaluate nine models…

Computation and Language · Computer Science 2026-04-24 Srishti Ginjala , Eric Fosler-Lussier , Christopher W. Myers , Srinivasan Parthasarathy

Relationships in scientific data, such as the numerical and spatial distribution relations of features in univariate data, the scalar-value combinations' relations in multivariate data, and the association of volumes in time-varying and…

Machine Learning · Computer Science 2022-07-25 Xiangyang He , Yubo Tao , Shuoliu Yang , Haoran Dai , Hai Lin

The cognitive reality of irregular morphological patterns has been debated for decades: do speakers extend them to novel forms, or are they lexical artifacts? A neural network trained on distributional input offers a learnability test: if…

Computation and Language · Computer Science 2026-05-26 Akhilesh Kakolu Ramarao , Kevin Tang , Dinah Baer-Henney

This paper investigates approximation-theoretic aspects of the in-context learning capability of the transformers in representing a family of noisy linear dynamical systems. Our first theoretical result establishes an upper bound on the…

Machine Learning · Computer Science 2025-10-22 Frank Cole , Yuxuan Zhao , Yulong Lu , Tianhao Zhang

Orthographic similarities across languages provide a strong signal for probabilistic decipherment, especially for closely related language pairs. The existing decipherment models, however, are not well-suited for exploiting these…

Computation and Language · Computer Science 2015-08-11 Iftekhar Naim , Daniel Gildea

Sign language recognition suffers from catastrophic scaling failure: models achieving high accuracy on small vocabularies collapse at realistic sizes. Existing architectures treat signs as atomic visual patterns, learning flat…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Bryan Cheng , Austin Jin , Jasper Zhang

We numerically investigate hyperuniformity in two-dimensional frictionless jammed packings of bidisperse systems. Hyperuniformity is characterized by the suppression of density fluctuations at large length scales, and the structure factor…

Soft Condensed Matter · Physics 2025-07-18 Duc T. Dam , Takeshi Kawasaki , Atsushi Ikeda , Kunimasa Miyazaki

Since their introduction, Transformer architectures have dominated Natural Language Processing (NLP). However, recent research has highlighted an inherent anisotropy phenomenon in these models, presenting a significant challenge to their…

Computation and Language · Computer Science 2026-04-13 Raphael Bernas , Fanny Jourdan , Antonin Poché , Céline Hudelot

Transformer-based table retrieval systems flatten structured tables into token sequences, making retrieval sensitive to the choice of serialization even when table semantics remain unchanged. We show that semantically equivalent…

Computation and Language · Computer Science 2026-04-29 Kushal Raj Bhandari , Adarsh Singh , Jianxi Gao , Soham Dan , Vivek Gupta

Transformers trained in low precision can suffer forward-error amplification. We give a first-order, module-wise theory that predicts when and where errors grow. For self-attention we derive a per-layer bound that factorizes into three…

Machine Learning · Computer Science 2025-10-28 Jinwoo Baek