English
Related papers

Related papers: US Code growth 1991-2025

200 papers

Conversations reflect the existing norms of a language. Previously, we found that utterance lengths in English fictional conversations in books and movies have shortened over a period of 200 years. In this work, we show that this shortening…

Physics and Society · Physics 2013-11-05 Christian M. Alis , May T. Lim

We present a comprehensive corpus of Russian primary and secondary legislation adopted between 1991 and 2025, comprising 304,382 texts (194,425,905 tokens). The corpus is available in two versions: the basic version contains texts with…

Computation and Language · Computer Science 2026-04-29 Denis Saveliev , Ruslan Kuchakov

This study is an attempt to build a contemporary linguistic corpus for Arabic language. The corpus produced, is a text corpus includes more than five million newspaper articles. It contains over a billion and a half words in total, out of…

Computation and Language · Computer Science 2016-11-15 Ibrahim Abu El-khair

This comprehensive survey examines the field of alphabetic codes, tracing their development from the 1960s to the present day. We explore classical alphabetic codes and their variants, analyzing their properties and the underlying…

Information Theory · Computer Science 2025-04-09 Roberto Bruno , Roberto De Prisco , Ugo Vaccaro

The problem of compression in standard information theory consists of assigning codes as short as possible to numbers. Here we consider the problem of optimal coding -- under an arbitrary coding scheme -- and show that it predicts Zipf's…

Computation and Language · Computer Science 2020-09-24 Ramon Ferrer-i-Cancho , Christian Bentz , Caio Seguin

We explore how ideas from infectious disease and genetics can be used to uncover patterns of cultural inheritance and innovation in a corpus of 591 national constitutions spanning 1789 - 2008. Legal "Ideas" are encoded as "topics" - words…

Social and Information Networks · Computer Science 2017-11-21 Daniel N. Rockmore , Chen Fang , Nicholas J. Foti , Tom Ginsburg , David C. Krakauer

For every $n\geq 27$, we show that the number of $n/(n-1)^+$-free words (i.e., threshold words) of length $k$ on $n$ letters grows exponentially in $k$. This settles all but finitely many cases of a conjecture of Ochem.

Combinatorics · Mathematics 2019-11-15 James D. Currie , Lucas Mol , Narad Rampersad

The mapping of lexical meanings to wordforms is a major feature of natural languages. While usage pressures might assign short words to frequent meanings (Zipf's law of abbreviation), the need for a productive and open-ended vocabulary,…

Computation and Language · Computer Science 2021-05-04 Tiago Pimentel , Irene Nikkarinen , Kyle Mahowald , Ryan Cotterell , Damián Blasi

Tandem duplication in DNA is the process of inserting a copy of a segment of DNA adjacent to the original position. Motivated by applications that store data in living organisms, Jain {\em et al.} (2016) proposed the study of codes that…

Combinatorics · Mathematics 2017-11-20 Yeow Meng Chee , Johan Chrisnata , Han Mao Kiah , Tuan Thanh Nguyen

Written language is a complex communication signal capable of conveying information encoded in the form of ordered sequences of words. Beyond the local order ruled by grammar, semantic and thematic structures affect long-range patterns in…

Physics and Society · Physics 2010-05-17 Marcelo A. Montemurro , Damian Zanette

We summarize the current state of the field of NLP & Law with a specific focus on recent technical and substantive developments. To support our analysis, we construct and analyze a nearly complete corpus of nearly one thousand NLP & Law…

Computation and Language · Computer Science 2026-05-13 Dirk Hartung , Daniel Martin Katz , Michael J. Bommarito , Lauritz Gerlach , Abhik Jana , Jerrold Soh

Large language models can now generate political messages as persuasive as those written by humans, raising concerns about how far this persuasiveness may continue to increase with model size. Here, we generate 720 persuasive messages on 10…

Computation and Language · Computer Science 2024-06-21 Kobi Hackenburg , Ben M. Tappin , Paul Röttger , Scott Hale , Jonathan Bright , Helen Margetts

Here I present an investigation on the evolution and use of vocabulary in data science in the last 13 years. Based on a rigorous statistical analysis, a database with 12,787 documents containing the words "data science" in the title,…

Digital Libraries · Computer Science 2022-04-22 Igor Barahona

We review recent progress in understanding the meaning of mutual information in natural language. Let us define words in a text as strings that occur sufficiently often. In a few previous papers, we have shown that a power-law distribution…

Information Theory · Computer Science 2020-03-11 Łukasz Dębowski

Extended variants of the recently introduced spread unary coding are described. These schemes, in which the length of the code word is fixed, allow representation of approximately n^2 numbers for n bits, rather than the n numbers of the…

Information Theory · Computer Science 2015-02-04 Subhash Kak

Scaling laws describe how language model capabilities grow with compute and data, but say nothing about how long a model matters once released. We provide the first large-scale empirical account of how scientists adopt and abandon language…

Digital Libraries · Computer Science 2026-04-10 Ana Trišović

We examine the complete dataset of baby name popularity collected by U.S. Social Security Administration for the last 131 years (1880-2010). The ranked baby name popularity can be fitted empirically by a piecewise function consisting of…

Adaptation and Self-Organizing Systems · Physics 2013-01-23 Wentian Li

Guess & Check (GC) codes are systematic binary codes that can correct multiple deletions, with high probability. GC codes have logarithmic redundancy in the length of the message $k$, and the encoding and decoding algorithms of these codes…

Information Theory · Computer Science 2019-05-01 Serge Kas Hanna , Salim El Rouayheb

Languages across the world exhibit Zipf's law of abbreviation, namely more frequent words tend to be shorter. The generalized version of the law - an inverse relationship between the frequency of a unit and its magnitude - holds also for…

Information Theory · Computer Science 2016-05-05 R. Ferrer-i-Cancho , C. Bentz , C. Seguin

Code Large Language Models (LLMs) are revolutionizing software engineering. However, scaling laws that guide the efficient training are predominantly analyzed on Natural Language (NL). Given the fundamental differences like strict syntax…

Computation and Language · Computer Science 2026-05-19 Xianzhen Luo , Wenzhen Zheng , Qingfu Zhu , Rongyi Zhang , Houyi Li , Siming Huang , YuanTao Fan , Wanxiang Che