English
Related papers

Related papers: Universal versus system-specific features of punct…

200 papers

Large-scale pretrained language models are the major driving force behind recent improvements in performance on the Winograd Schema Challenge, a widely employed test of common sense reasoning ability. We show, however, with a new diagnostic…

Computation and Language · Computer Science 2020-05-08 Mostafa Abdou , Vinit Ravishankar , Maria Barrett , Yonatan Belinkov , Desmond Elliott , Anders Søgaard

The uniform information density (UID) hypothesis proposes that speakers aim to distribute information evenly throughout a text, balancing production effort and listener comprehension difficulty. However, language typically does not maintain…

Computation and Language · Computer Science 2025-06-05 Eleftheria Tsipidi , Samuel Kiegeland , Franz Nowak , Tianyang Xu , Ethan Wilcox , Alex Warstadt , Ryan Cotterell , Mario Giulianelli

The limited range in its abscissa of ranked letter frequency distributions causes multiple functions to fit the observed distribution reasonably well. In order to critically compare various functions, we apply the statistical model…

Computation and Language · Computer Science 2012-05-07 Wentian Li , Pedro Miramontes

A standard measure of the influence of a research paper is the number of times it is cited. However, papers may be cited for many reasons, and citation count offers limited information about the extent to which a paper affected the content…

Computation and Language · Computer Science 2022-10-26 Sandeep Soni , David Bamman , Jacob Eisenstein

A word~$w$ has a border $u$ if $u$ is a non-empty proper prefix and suffix of $u$. A word~$w$ is said to be \emph{closed} if $w$ is of length at most $1$ or if $w$ has a border that occurs exactly twice in $w$. A word~$w$ is said to be…

Combinatorics · Mathematics 2024-05-24 Daniel Gabric

Individual happiness is a fundamental societal metric. Normally measured through self-report, happiness has often been indirectly characterized and overshadowed by more readily quantifiable economic indicators such as gross domestic…

We consider the spreading and competition of languages that are spoken by a population of individuals. The individuals can change their mother tongue during their lifespan, pass on their language to their offspring and finally die. The…

Physics and Society · Physics 2009-11-11 Tiberiu Teşileanu , Hildegard Meyer-Ortmanns

With the advancement of information systems, means of communications are becoming cheaper, faster and more available. Today, millions of people carrying smart-phones or tablets are able to communicate at practically any time and anywhere…

Social and Information Networks · Computer Science 2021-09-02 Pedro O. S. Vaz de Melo , Christos Faloutsos , Renato Assunção , Rodrigo Alves , Antonio A. F. Loureiro

Driven by recent advances AI, we passengers are entering a golden age of scientific discovery. But golden for whom? Confronting our insecurity that others may beat us to the most acclaimed breakthroughs of the era, we propose a novel…

Digital Libraries · Computer Science 2023-04-04 Samuel Albanie , Liliane Momeni , João F. Henriques

Language enables humans to share knowledge, reason about the world, and pass on strategies for survival and innovation across generations. At the heart of this process is not just the ability to communicate but also the remarkable…

Computation and Language · Computer Science 2026-02-25 Jan Philip Wahle

Language is a uniquely human trait, conveying information efficiently by organizing word sequences in sentences into hierarchical structures. A central question persists: Why is human language hierarchical? In this study, we show that…

Computation and Language · Computer Science 2026-01-07 Luyao Chen , Weibo Gao , Junjie Wu , Jinshan Wu , Angela D. Friederici

Language similarities can be caused by genetic relatedness, areal contact, universality, or chance. Colexification, i.e. a type of similarity where a single lexical form is used to convey multiple meanings, is underexplored. In our work, we…

Computation and Language · Computer Science 2024-01-08 Yiyi Chen , Johannes Bjerva

This paper introduces a new generalization of the power generalized Weibull distribution called the generalized power generalized Weibull distribution. This distribution can also be considered as a generalization of Weibull distribution.…

Statistics Theory · Mathematics 2018-10-16 Mahmoud Ali Selim

We present in this paper a numerical investigation of literary texts by various well-known English writers, covering the first half of the twentieth century, based upon the results obtained through corpus analysis of the texts. A fractal…

Other Condensed Matter · Physics 2009-11-11 L. L. Goncalves , L. B. Goncalves

People tend to distribute information evenly in language production for better and clearer communication. In this study, we compared essays written by second language learners with various native language (L1) backgrounds to investigate how…

Computation and Language · Computer Science 2024-11-07 Zixin Tang , Janet G. van Hell

We study the cycle structure of words in several random permutations. We assume that the permutations are independent and that their distribution is conjugation invariant, with a good control on their short cycles. If, after successive…

Combinatorics · Mathematics 2023-10-24 Mohamed Slim Kammoun , Mylène Maïda

A word $u$ is a scattered factor of $w$ if $u$ can be obtained from $w$ by deleting some of its letters. That is, there exist the (potentially empty) words $u_1,u_2,..., u_n$, and $v_0,v_1,..,v_n$ such that $u = u_1u_2...u_n$ and $w =…

Formal Languages and Automata Theory · Computer Science 2019-05-27 Joel D. Day , Pamela Fleischmann , Florin Manea , Dirk Nowotka

Many studies were recently done for investigating the properties of contextual language models but surprisingly, only a few of them consider the properties of these models in terms of semantic similarity. In this article, we first focus on…

Computation and Language · Computer Science 2021-11-25 Olivier Ferret

Languages across the world exhibit Zipf's law of abbreviation, namely more frequent words tend to be shorter. The generalized version of the law - an inverse relationship between the frequency of a unit and its magnitude - holds also for…

Information Theory · Computer Science 2016-05-05 R. Ferrer-i-Cancho , C. Bentz , C. Seguin

We study the frequency distributions and correlations of the word lengths of ten European languages. Our findings indicate that a) the word-length distribution of short words quantified by the mean value and the entropy distinguishes the…