English
Related papers

Related papers: Two halves of a meaningful text are statistically …

200 papers

Given two candidate models, and a set of target observations, we address the problem of measuring the relative goodness of fit of the two models. We propose two new statistical tests which are nonparametric, computationally efficient…

There are two methods for counting the number of occurrences of a string in another large string. One is to count the number of places where the string is found. The other is to determine how many pieces of string can be extracted without…

Data Structures and Algorithms · Computer Science 2022-11-09 Ayaka Takamoto , Mitsuo Yoshida , Kyoji Umemura

Textual geographic information is indispensable and heavily relied upon in practical applications. The absence of clear distribution poses challenges in effectively harnessing geographic information, thereby driving our quest for…

Computation and Language · Computer Science 2023-09-04 Zhenhua Wang , Daiyu Zhang , Ming Ren , Guang Xu

We develop a new rank-based approach for univariate two-sample testing in the presence of missing data which makes no assumptions about the missingness mechanism. This approach is a theoretical extension of the Wilcoxon-Mann-Whitney test…

Methodology · Statistics 2024-03-25 Yijin Zeng , Niall M. Adams , Dean A. Bodenham

As the context length that large language models can handle continues to increase, these models demonstrate an enhanced ability to utilize distant information for tasks such as language modeling. This capability contrasts with human reading…

Computation and Language · Computer Science 2024-06-18 Yutong Hu , Quzhe Huang , Kangcheng Luo , Yansong Feng

Lexical Semantic Change (LSC) is the phenomenon in which the meaning of a word change over time. Most studies on LSC focus on improving the performance of estimating the degree of LSC, however, it is often difficult to interpret how the…

Computation and Language · Computer Science 2026-02-11 Kohei Oda , Hiroya Takamura , Kiyoaki Shirai , Natthawut Kertkeidkachorn

Many networks in natural and human-made systems exhibit scale-free properties and are small worlds. Now we show that people's understanding of complex systems in their cognitive maps also follow a scale-free topology (P_k = k^-lambda,…

Neurons and Cognition · Quantitative Biology 2007-05-23 Uygar Ozesmi , Can Ozan Tan

Most research on natural language processing treats bias as an absolute concept: Based on a (probably complex) algorithmic analysis, a sentence, an article, or a text is classified as biased or not. Given the fact that for humans the…

Computation and Language · Computer Science 2022-10-14 Alonso Palomino , Martin Potthast , Khalid Al-Khatib , Benno Stein

The use of content features, particularly textual and linguistic for fake news detection is under-researched, despite empirical evidence showing the features could contribute to differentiating real and fake news. To this end, this study…

Computation and Language · Computer Science 2026-05-11 Vimala Balakrishnan , Lee Zing Hii , Eric Laporte

Why do some things succeed in the marketplace of ideas? While some argue that content drives success, others suggest that style, or the way ideas are presented, also plays an important role. To provide a stringent test of style's…

Computation and Language · Computer Science 2022-01-11 Reihane Boghrati , Jonah Berger , Grant Packard

We use the formulation of equilibrium statistical mechanics in order to study some important characteristics of language. Using a simple expression for the Hamiltonian of a language system, which is directly implied by the Zipf law, we are…

Physics and Society · Physics 2009-11-11 Kosmas Kosmidis , Alkiviadis Kalampokis , Panos Argyrakis

We show that the mutual information between two symbols, as a function of the number of symbols between the two, decays exponentially in any probabilistic regular grammar, but can decay like a power law for a context-free grammar. This…

Disordered Systems and Neural Networks · Physics 2017-08-25 Henry W. Lin , Max Tegmark

Much of statistics relies upon four key elements: a law of large numbers, a calculus to operationalize stochastic convergence, a central limit theorem, and a framework for constructing local approximations. These elements are…

Optimization and Control · Mathematics 2018-01-09 Anil Aswani

The main goal is to develop and, consequently, compare stochastic methods for detection whether a structural change in panel data occurred at some unknown time or not. Panel data of our interest consist of a moderate or relatively large…

Methodology · Statistics 2016-08-22 Barbora Peštová , Michal Pešta

A statistical physics study of punctuation effects on sentence lengths is presented for written texts: {\it Alice in wonderland} and {\it Through a looking glass}. The translation of the first text into esperanto is also considered as a…

Computation and Language · Computer Science 2012-09-04 M. Ausloos

The explanations of large language models have recently been shown to be sensitive to the randomness used for their training, creating a need to characterize this sensitivity. In this paper, we propose a characterization that questions the…

Computation and Language · Computer Science 2024-03-18 Jeremie Bogaert , Francois-Xavier Standaert

We study the local limit distribution of the number of occurrences of a symbol in words of length $n$ generated at random in a regular language according to a rational stochastic model. We present an analysis of the main local limits when…

Probability · Mathematics 2021-02-19 Massimiliano Goldwurm , Jianyi Lin , Marco Vignati

Computing the {\em matching statistics} of a string $P[1..m]$ with respect to a text $T[1..n]$ is a fundamental problem which has application to genome sequence comparison. In this paper, we study the problem of computing the matching…

Data Structures and Algorithms · Computer Science 2022-01-14 Younan Gao

Written language is complex. A written text can be considered an attempt to convey a meaningful message which ends up being constrained by language rules, context dependence and highly redundant in its use of resources. Despite all these…

Computation and Language · Computer Science 2019-05-20 E. Estevez-Rams , A. Mesa Rodriguez , D. Estevez-Moya

This paper offers a commentary on the use of notions of statistical significance in choice modelling. We review the reasons for uncertainty in parameter estimates, provide a precise discussion on the computation of measures of uncertainty…

Econometrics · Economics 2026-05-18 Stephane Hess , Andrew Daly , Michiel Bliemer , Angelo Guevara , Ricardo Daziano , Thijs Dekker