English
Related papers

Related papers: In narrative texts punctuation marks obey the same…

200 papers

Long-range correlations are found in symbolic sequences from human language, music and DNA. Determining the span of correlations in dolphin whistle sequences is crucial for shedding light on their communicative complexity. Dolphin whistles…

Neurons and Cognition · Quantitative Biology 2014-12-03 Ramon Ferrer-i-Cancho , Brenda McCowan

Labeling of sentence boundaries is a necessary prerequisite for many natural language processing tasks, including part-of-speech tagging and sentence alignment. End-of-sentence punctuation marks are ambiguous; to disambiguate them most…

cmp-lg · Computer Science 2008-02-03 David D. Palmer , Marti A. Hearst

Recently, it has been claimed that a linear relationship between a measure of information content and word length is expected from word length optimization and it has been shown that this linearity is supported by a strong correlation…

Data Analysis, Statistics and Probability · Physics 2019-12-11 Ramon Ferrer-i-Cancho , Fermín Moscoso del Prado Martín

The pioneering research of G. K. Zipf on the relationship between word frequency and other word features led to the formulation of various linguistic laws. The most popular is Zipf's law for word frequencies. Here we focus on two laws that…

Computation and Language · Computer Science 2020-09-24 Bernardino Casas , Antoni Hernández-Fernández , Neus Català , Ramon Ferrer-i-Cancho , Jaume Baixeries

We investigate the origin of Zipf's law for words in written texts by means of a stochastic dynamical model for text generation. The model incorporates both features related to the general structure of languages and memory effects inherent…

Statistical Mechanics · Physics 2007-05-23 Damián H. Zanette , Marcelo A. Montemurro

Background: Zipf's discovery that word frequency distributions obey a power law established parallels between biological and physical processes, and language, laying the groundwork for a complex systems perspective on human communication.…

Computation and Language · Computer Science 2009-11-11 Eduardo G. Altmann , Janet B. Pierrehumbert , Adilson E. Motter

The dependence with text length of the statistical properties of word occurrences has long been considered a severe limitation quantitative linguistics. We propose a simple scaling form for the distribution of absolute word frequencies…

Physics and Society · Physics 2015-06-15 Francesc Font-Clos , Gemma Boleda , Álvaro Corral

Natural languages are full of rules and exceptions. One of the most famous quantitative rules is Zipf's law which states that the frequency of occurrence of a word is approximately inversely proportional to its rank. Though this `law' of…

Computation and Language · Computer Science 2015-05-27 Jake Ryland Williams , James P. Bagrow , Christopher M. Danforth , Peter Sheridan Dodds

The Zipf's law is the major regularity of statistical linguistics that served as a prototype for rank-frequency relations and scaling laws in natural sciences. Here we show that the Zipf's law -- together with its applicability for a single…

Data Analysis, Statistics and Probability · Physics 2015-06-15 Armen E. Allahverdyan , Weibing Deng , Q. A. Wang

Over the past two centuries, the frequency of word usage in major Western languages has exhibited small amplitude regular cycles, superimposed on larger background trends. We show that these cycles of word usage organize into semantically…

Physics and Society · Physics 2026-05-01 Alejandro Pardo Pintos , Diego E Shalom , Guillermo Cecchi , Gabriel Mindlin , Marcos A Trevisan

Punctuation plays a vital role in structuring meaning, yet current models often struggle to restore it accurately in transcripts of spontaneous speech, especially in the presence of disfluencies such as false starts and backtracking. These…

Computation and Language · Computer Science 2025-06-05 Sidharth Pulipaka , Sparsh Jain , Ashwin Sankar , Raj Dabre

Probabilistic approaches to part-of-speech tagging rely primarily on whole-word statistics about word/tag combinations as well as contextual information. But experience shows about 4 per cent of tokens encountered in test sets are unknown…

Computation and Language · Computer Science 2013-02-28 Greg Adams , Beth Millar , Eric Neufeld , Tim Philip

The article presents a new interpretation for Zipf-Mandelbrot's law in natural language which rests on two areas of information theory. Firstly, we construct a new class of grammar-based codes and, secondly, we investigate properties of…

Information Theory · Computer Science 2020-03-11 Łukasz Dębowski

In this paper we explore where information is collected and how it is propagated throughout layers in large language models (LLMs). We begin by examining the surprising computational importance of punctuation tokens which previous work has…

Computation and Language · Computer Science 2025-08-21 Sonakshi Chauhan , Maheep Chaudhary , Koby Choy , Samuel Nellessen , Nandi Schoots

We describe a statistical approach for modeling dialogue acts in conversational speech, i.e., speech-act-like units such as Statement, Question, Backchannel, Agreement, Disagreement, and Apology. Our model detects and predicts dialogue acts…

Computation and Language · Computer Science 2022-02-28 A. Stolcke , K. Ries , N. Coccaro , E. Shriberg , R. Bates , D. Jurafsky , P. Taylor , R. Martin , C. Van Ess-Dykema , M. Meteer

Words are fundamental linguistic units that connect thoughts and things through meaning. However, words do not appear independently in a text sequence. The existence of syntactic rules induces correlations among neighboring words. Using an…

Computation and Language · Computer Science 2023-03-15 David Sanchez , Luciano Zunino , Juan De Gregorio , Raul Toral , Claudio Mirasso

Neural network-based embeddings have been the mainstream approach for creating a vector representation of the text to capture lexical and semantic similarities and dissimilarities. In general, existing encoding methods dismiss the…

Computation and Language · Computer Science 2022-08-09 Mansooreh Karami , Ahmadreza Mosallanezhad , Michelle V Mancenido , Huan Liu

We explore a probabilistic model of an artistic text: words of the text are chosen independently of each other in accordance with a discrete probability distribution on an infinite dictionary. The words are enumerated 1, 2, $\ldots$, and…

Statistics Theory · Mathematics 2019-05-02 Mikhail Chebunin , Artyom Kovalevskii

We demonstrate that large texts, representing human (English, Russian, Ukrainian) and artificial (C++, Java) languages, display quantitative patterns characterized by the Benford-like and Zipf laws. The frequency of a word following the…

Computation and Language · Computer Science 2018-03-13 Evgeny Shulzinger , Irina Legchenkova , Edward Bormashenko

We describe an approach to robust domain-independent syntactic parsing of unrestricted naturally-occurring (English) input. The technique involves parsing sequences of part-of-speech and punctuation labels using a unification-based grammar…

cmp-lg · Computer Science 2008-02-03 Ted Briscoe , John Carroll