English
Related papers

Related papers: Two halves of a meaningful text are statistically …

200 papers

Divergent word usages reflect differences among people. In this paper, we present a novel angle for studying word usage divergence -- word interpretations. We propose an approach that quantifies semantic differences in interpretations among…

Computation and Language · Computer Science 2017-03-30 Tianran Hu , Ruihua Song , Maya Abtahian , Philip Ding , Xing Xie , Jiebo Luo

Many empirical time series are genuinely symbolic: examples range from link activation patterns in network science, DNA coding or firing patterns in neuroscience to cryptography or combinatorics on words. In some other contexts, the…

Chaotic Dynamics · Physics 2023-07-19 Lluis Arola-Fernandez , Lucas Lacasa

Statistical methods have been widely employed in recent years to grasp many language properties. The application of such techniques have allowed an improvement of several linguistic applications, which encompasses machine translation,…

Computation and Language · Computer Science 2016-02-22 Henrique F. de Arruda , Luciano da F. Costa , Diego R. Amancio

Zipf's law is just one out of many universal laws proposed to describe statistical regularities in language. Here we review and critically discuss how these laws can be statistically interpreted, fitted, and tested (falsified). The modern…

Physics and Society · Physics 2016-05-27 Eduardo G. Altmann , Martin Gerlach

The surge in digitized text data requires reliable inferential methods on observed textual patterns. This article proposes a novel two-sample text test for comparing similarity between two groups of documents. The hypothesis is whether the…

Machine Learning · Statistics 2025-05-09 Jingbin Xu , Chen Qian , Meimei Liu , Feng Guo

Words shift in meaning for many reasons, including cultural factors like new technologies and regular linguistic processes like subjectification. Understanding the evolution of language and culture requires disentangling these underlying…

Computation and Language · Computer Science 2016-09-27 William L. Hamilton , Jure Leskovec , Dan Jurafsky

The distribution of frequency counts of distinct words by length in a language's vocabulary will be analyzed using two methods. The first, will look at the empirical distributions of several languages and derive a distribution that…

Computation and Language · Computer Science 2012-07-17 Reginald D. Smith

Comparison of statistical models (experiments) is an important branch of mathematical statistics, which gives deep insights in many aspects of foundation of statistics. So far, there are two quantum versions of the concept: Comparison with…

Quantum Physics · Physics 2014-09-22 Keiji Matsumoto

Understanding texts requires memory: the reader has to keep in mind enough words to create meaning. This calls for a relation between the memory of the reader and the structure of the text. To investigate this interaction, we first identify…

Physics and Society · Physics 2007-05-23 E. Alvarez-Lacalle , B. Dorow , J. -P. Eckmann , E. Moses

This paper compares a qualitative reasoning model of translation with a quantitative statistical model. We consider these models within the context of two hypothetical speech translation systems, starting with a logic-based design and…

cmp-lg · Computer Science 2008-02-03 Hiyan Alshawi

We evaluated the impact of changing the observation scale over the entropy measures for text descriptions. MIDI coded Music, computer code and two human natural languages were studied at the scale of characters, words, and at the…

Information Theory · Computer Science 2017-01-13 Gerardo Febres , Klaus Jaffe

Knowing which strings in a massive text are significant -- that is, which strings are common and distinct from other strings -- is valuable for several applications, including text compression and tokenization. Frequency in itself is not…

Data Structures and Algorithms · Computer Science 2024-04-24 Peaker Guo , Patrick Eades , Anthony Wirth , Justin Zobel

For testing the statistical significance of a treatment effect, we usually compare between two parts of a population, one is exposed to the treatment, and the other is not exposed to it. Standard parametric and nonparametric two-sample…

Computation · Statistics 2012-11-02 Bikram Karmakar , Kumaresh Dhara , Kushal Kumar Dey , Analabha Basu , Anil Ghosh

We consider multivariate two-sample tests of means, where the location shift between the two populations is expected to be related to a known graph structure. An important application of such tests is the detection of differentially…

Quantitative Methods · Quantitative Biology 2014-05-16 Laurent Jacob , Pierre Neuvial , Sandrine Dudoit

Understanding semantic relations between two texts is crucial for many information and document management tasks, in which one must determine whether the content fully overlaps, is completely superseded by another document, or overlaps only…

Computation and Language · Computer Science 2025-12-02 Yehudit Aperstein , Alon Gottlib , Gal Benita , Alexander Apartsin

Measuring similarity is a basic task in information retrieval, and now often a building-block for more complex arguments about cultural change. But do measures of textual similarity and distance really correspond to evidence about cultural…

Computation and Language · Computer Science 2018-07-03 Ted Underwood

With limited resources, scientific inquiries must be prioritized for further study, funding, and translation based on their practical significance: whether the effect size is large enough to be meaningful in the real world. Doing so must…

Methodology · Statistics 2022-05-27 Bruce A. Corliss , Yaotian Wang , Heman Shakeri , Philip E. Bourne

The basic properties of the Fisher information allow to reveal the statistical meaning of classical inequalities between mean functions. The properties applied to scale mixtures of Gaussian distributions lead to a new mean function of…

Statistics Theory · Mathematics 2019-04-09 Abram M. Kagan , Paul J. Smith

The controversy about statistical significance vs. scientific relevance is more than 100 years old. But still nowadays null hypothesis significance testing is considered as gold standard in many empirical fields from economics and social…

Applications · Statistics 2022-11-23 Uwe Hassler

In natural language using short sentences is considered efficient for communication. However, a text composed exclusively of such sentences looks technical and reads boring. A text composed of long ones, on the other hand, demands…