English
Related papers

Related papers: Verifying Heaps' law using Google Books Ngram data

200 papers

Word embeddings are a powerful approach for analyzing language, and exponential family embeddings (EFE) extend them to other types of data. Here we develop structured exponential family embeddings (S-EFE), a method for discovering…

Computation and Language · Computer Science 2017-10-03 Maja Rudolph , Francisco Ruiz , Susan Athey , David Blei

Natural language exhibits statistical dependencies at a wide range of scales. For instance, the mutual information between words in natural language decays like a power law with the temporal lag between them. However, many statistical…

Computation and Language · Computer Science 2019-12-17 Aakash Sarkar , Marc Howard

Introduction: Clinical text classification using natural language processing (NLP) models requires adequate training data to achieve optimal performance. For that, 200-500 documents are typically annotated. The number is constrained by time…

Computation and Language · Computer Science 2026-01-23 Jaya Chaturvedi , Saniya Deshpande , Chenkai Ma , Robert Cobb , Angus Roberts , Robert Stewart , Daniel Stahl , Diana Shamsutdinova

Here we test Neutral models against the evolution of English word frequency and vocabulary at the population scale, as recorded in annual word frequencies from three centuries of English language books. Against these data, we test both…

Computation and Language · Computer Science 2017-04-03 Damian Ruck , R. Alexander Bentley , Alberto Acerbi , Philip Garnett , Daniel J. Hruschka

We present in this paper a numerical investigation of literary texts by various well-known English writers, covering the first half of the twentieth century, based upon the results obtained through corpus analysis of the texts. A fractal…

Other Condensed Matter · Physics 2009-11-11 L. L. Goncalves , L. B. Goncalves

The syntactic structure of a sentence can be represented as a graph, where vertices are words and edges indicate syntactic dependencies between them. In this setting, the distance between two linked words is defined as the difference…

Computation and Language · Computer Science 2025-08-12 Sonia Petrini , Ramon Ferrer-i-Cancho

This paper investigates the information encoded in the embeddings of large language models (LLMs). We conduct simulations to analyze the representation entropy and discover a power law relationship with model sizes. Building upon this…

Machine Learning · Computer Science 2024-02-07 Zhiquan Tan , Chenghai Li , Weiran Huang

Lexical ambiguity is widespread in language, allowing for the reuse of economical word forms and therefore making language more efficient. If ambiguous words cannot be disambiguated from context, however, this gain in efficiency might make…

Computation and Language · Computer Science 2024-05-29 Tiago Pimentel , Rowan Hall Maudslay , Damián Blasi , Ryan Cotterell

We analyse the dependence of stock return cross-correlations on the sampling frequency of the data known as the Epps effect: For high resolution data the cross-correlations are significantly smaller than their asymptotic value as observed…

Statistical Finance · Quantitative Finance 2009-10-26 Bence Toth , Janos Kertesz

We investigate symbolic sequences and in particular information carriers as e.g. books and DNA-strings. First the higher order Shannon entropies are calculated, a characteristic root law is detected. Then the algorithmic entropy is…

Disordered Systems and Neural Networks · Physics 2007-05-23 Werner Ebeling , Alexander Neiman , Thorsten Poeschel

We consider probabilistic topic models and more recent word embedding techniques from a perspective of learning hidden semantic representations. Inspired by a striking similarity of the two approaches, we merge them and learn probabilistic…

Computation and Language · Computer Science 2017-11-15 Anna Potapenko , Artem Popov , Konstantin Vorontsov

Semantic type mismatch between a noun and its context is central to coercion phenomena. This paper introduces a graph-based method to examine how lexical and contextual type information is reflected in word embeddings. We select nouns from…

Computation and Language · Computer Science 2026-05-25 Long Chen , Deniz Ekin Yavas

Can self-organization of scientific communication be specified by using literature-based indicators? In this study, we explore this question by applying entropy measures to typical "Mode-2" fields of knowledge production. We hypothesized…

Information Retrieval · Computer Science 2010-01-11 Gaston Heimeriks , Loet Leydesdorff , Peter Van den Besselaar

It has been shown in a recent publication that words in human-produced English language tend to have an information content close to the conditional entropy. In this paper, we show that the same is true for events in human-produced…

Sound · Computer Science 2022-11-24 Mathias Rose Bjare , Stefan Lattner

It is now a common practice to compare models of human language processing by predicting participant reactions (such as reading times) to corpora consisting of rich naturalistic linguistic materials. However, many of the corpora used in…

Computation and Language · Computer Science 2017-08-22 Richard Futrell , Edward Gibson , Hal Tily , Idan Blank , Anastasia Vishnevetsky , Steven T. Piantadosi , Evelina Fedorenko

Social bias in language - towards genders, ethnicities, ages, and other social groups - poses a problem with ethical impact for many NLP applications. Recent research has shown that machine learning models trained on respective data may not…

Computation and Language · Computer Science 2020-11-25 Maximilian Spliethöver , Henning Wachsmuth

How much data is required to learn the structure of a language via next-token prediction? We study this question for synthetic datasets generated via a Probabilistic Context-Free Grammar (PCFG) -- a tree-like generative model that captures…

Computation and Language · Computer Science 2024-10-30 Francesco Cagnetta , Matthieu Wyart

We present the new empirical parameter $f_c$, the most probable usage frequency of a word in a language, computed via the distribution of documents over frequency $x$ of the word. This parameter allows for filtering the core lexicon of a…

Disordered Systems and Neural Networks · Physics 2007-05-23 Dmitri Volchenkov , Philippe Blanchard , Serge Sharoff

Recent work has found that contemporary language models such as transformers can become so good at next-word prediction that the probabilities they calculate become worse for predicting reading time. In this paper, we propose that this can…

Computation and Language · Computer Science 2026-03-11 James A. Michaelov , Roger P. Levy

As is the case of many signals produced by complex systems, language presents a statistical structure that is balanced between order and disorder. Here we review and extend recent results from quantitative characterisations of the degree of…

Computation and Language · Computer Science 2015-03-05 Marcelo A Montemurro , Damián H Zanette
‹ Prev 1 8 9 10 Next ›