English
Related papers

Related papers: Correlation Dimension of Natural Language in a Sta…

200 papers

Large language models (LLMs) have complicated internal dynamics, but induce representations of words and phrases whose geometry we can study. Human language processing is also opaque, but neural response measurements can provide (noisy)…

Computation and Language · Computer Science 2023-11-01 Jiaang Li , Antonia Karamolegkou , Yova Kementchedjhieva , Mostafa Abdou , Sune Lehmann , Anders Søgaard

Calculating the semantic similarity between sentences is a long dealt problem in the area of natural language processing. The semantic analysis field has a crucial role to play in the research related to the text analytics. The semantic…

Computation and Language · Computer Science 2018-02-22 Atish Pawar , Vijay Mago

This paper introduces the correlation-of-divergency coefficient, c-delta, a custom statistical measure designed to quantify the similarity of internal divergence patterns between two groups of values. Unlike conventional correlation…

Methodology · Statistics 2026-03-10 Johan F. Hoorn

We study adaptive data-dependent dimensionality reduction in the context of supervised learning in general metric spaces. Our main statistical contribution is a generalization bound for Lipschitz functions in metric spaces that are…

Machine Learning · Computer Science 2015-03-26 Lee-Ad Gottlieb , Aryeh Kontorovich , Robert Krauthgamer

Calculating semantic textual similarity is a foundational task in natural language processing. Current large language models (LLMs) based methods typically rely on extracting last-layer hidden states with fixed dimensions to compute…

Computation and Language · Computer Science 2026-05-29 Kaijie Zheng , Weiqin Wang , Yile Wang , Hui Huang

We study the fractal structure of language, aiming to provide a precise formalism for quantifying properties that may have been previously suspected but not formally shown. We establish that language is: (1) self-similar, exhibiting…

Computation and Language · Computer Science 2024-05-24 Ibrahim Alabdulmohsin , Vinh Q. Tran , Mostafa Dehghani

This paper proposes a simple test for compositionality (i.e., literal usage) of a word or phrase in a context-specific way. The test is computationally simple, relying on no external resources and only uses a set of trained word vectors.…

Computation and Language · Computer Science 2016-11-30 Hongyu Gong , Suma Bhat , Pramod Viswanath

Dimensional analysis (DA) pays attention to fundamental physical dimensions such as length and mass when modelling scientific and engineering systems. It goes back at least a century to Buckingham's Pi theorem, which characterizes a…

Machine Learning · Computer Science 2023-12-19 G. Alexi Rodriguez-Arelis , William J. Welch

Audio Large Language Models (Audio LLMs) have demonstrated strong capabilities in integrating speech perception with language understanding. However, whether their internal representations align with human neural dynamics during…

Sound · Computer Science 2026-02-04 Haoyun Yang , Xin Xiao , Jiang Zhong , Yu Tian , Dong Xiaohua , Yu Mao , Hao Wu , Kaiwen Wei

Metrics for measuring the comparability of corpora or texts need to be developed and evaluated systematically. Applications based on a corpus, such as training Statistical MT systems in specialised narrow domains, require finding a…

Computation and Language · Computer Science 2014-04-16 Bogdan Babych , Anthony Hartley

The review summarizes the main methodological concepts used in studying natural language from the perspective of complexity science and documents their applicability in identifying both universal and system-specific features of language in…

Physics and Society · Physics 2024-01-09 Tomasz Stanisz , Stanisław Drożdż , Jarosław Kwapień

The class of stochastically self-similar sets contains many famous examples of random sets, e.g. Mandelbrot percolation and general fractal percolation. Under the assumption of the uniform open set condition and some mild assumptions on the…

Metric Geometry · Mathematics 2019-12-23 Sascha Troscheit

Closely related languages show linguistic similarities that allow speakers of one language to understand speakers of another language without having actively learned it. Mutual intelligibility varies in degree and is typically tested in…

Computation and Language · Computer Science 2024-02-06 Jessica Nieder , Johann-Mattis List

The rapid development of such natural language processing tasks as style transfer, paraphrase, and machine translation often calls for the use of semantic similarity metrics. In recent years a lot of methods to measure the semantic…

Computation and Language · Computer Science 2022-11-15 Ivan P. Yamshchikov , Viacheslav Shibaev , Nikolay Khlebnikov , Alexey Tikhonov

Given a random text over a finite alphabet, we study the frequencies at which fixed-length words occur as subsequences. As the data size grows, the joint distribution of word counts exhibits a rich asymptotic structure. We investigate all…

Probability · Mathematics 2026-05-06 Chaim Even-Zohar , Tsviqa Lakrec , Ran J. Tessler

Scientific discovery increasingly depends on efficient experimental optimization to navigate vast design spaces under time and resource constraints. Traditional approaches often require extensive domain expertise and feature engineering.…

Machine Learning · Computer Science 2025-11-10 Bojana Ranković , Ryan-Rhys Griffiths , Philippe Schwaller

A constructive version of Hausdorff dimension is developed using constructive supergales, which are betting strategies that generalize the constructive supermartingales used in the theory of individual random sequences. This constructive…

Computational Complexity · Computer Science 2007-05-23 Jack H. Lutz

A measure of similarity between text embeddings can be considered adequate only if it adheres to the human perception of similarity between texts. In this paper, we introduce the distance-to-distance ratio (DDR), a novel measure of…

Computation and Language · Computer Science 2026-01-27 Abdullah Qureshi , Kenneth Rice , Alexander Wolpert

Measuring the similarity of short written contexts is a fundamental problem in Natural Language Processing. This article provides a unifying framework by which short context problems can be categorized both by their intended application and…

Computation and Language · Computer Science 2010-10-19 Ted Pedersen

In this article we propose a novel method to estimate the frequency distribution of linguistic variables while controlling for statistical non-independence due to shared ancestry. Unlike previous approaches, our technique uses all available…

Populations and Evolution · Quantitative Biology 2021-03-22 Gerhard Jäger , Johannes Wahle
‹ Prev 1 4 5 6 7 8 10 Next ›