English
Related papers

Related papers: Maximum Entropy, Word-Frequency, Chinese Character…

200 papers

Shannon entropy is often a quantity of interest to linguists studying the communicative capacity of human language. However, entropy must typically be estimated from observed data because researchers do not have access to the underlying…

Computation and Language · Computer Science 2022-04-06 Aryaman Arora , Clara Meister , Ryan Cotterell

The task of finding a criterion allowing to distinguish a text from an arbitrary set of words is rather relevant in itself, for instance, in the aspect of development of means for internet-content indexing or separating signals and noise in…

Computation and Language · Computer Science 2007-10-02 D. V. Lande , A. A. Snarskii

Let $W^{(n)}$ be the $n$-letter word obtained by repeating a fixed word $W$, and let $R_n$ be a random $n$-letter word over the same alphabet. We show several results about the length of the longest common subsequence (LCS) between…

Probability · Mathematics 2021-06-07 Boris Bukh , Christopher Cox

By determining which were the most common English words and phrases since the beginning of the 16th century, we obtain a unique large-scale view of the evolution of written text. We find that the most common words and phrases in any given…

Physics and Society · Physics 2012-12-10 Matjaz Perc

The principle of maximum entropy is a broadly applicable technique for computing a distribution with the least amount of information possible while constrained to match empirically estimated feature expectations. However, in many real-world…

Machine Learning · Computer Science 2022-08-16 Kenneth Bogert , Yikang Gui , Prashant Doshi

We investigate the behavior of the periods and border lengths of random words over a fixed alphabet. We show that the asymptotic probability that a random word has a given maximal border length $k$ is a constant, depending only on $k$ and…

Formal Languages and Automata Theory · Computer Science 2019-12-18 Štěpán Holub , Jeffrey Shallit

Words are sequences of letters over a finite alphabet. We study two intimately related topics for this object: quasi-randomness and limit theory. With respect to the first topic we investigate the notion of uniform distribution of letters…

Combinatorics · Mathematics 2021-09-01 Hiêp Hàn , Marcos Kiwi , Matías Pavez-Signé

Quantitative linguistics has provided us with a number of empirical laws that characterise the evolution of languages and competition amongst them. In terms of language usage, one of the most influential results is Zipf's law of word…

Physics and Society · Physics 2009-01-21 Alvaro Corral , Ramon Ferrer-i-Cancho , Gemma Boleda , Albert Diaz-Guilera , .

In this paper, we propose a joint algorithm for the word segmentation on Chinese social media. Previous work mainly focus on word segmentation for plain Chinese text, in order to develop a Chinese social media processing tool, we need to…

Computation and Language · Computer Science 2015-10-27 Yao Yushi , Huang Zheng

In this paper, we propose new methods to learn Chinese word representations. Chinese characters are composed of graphical components, which carry rich semantics. It is common for a Chinese learner to comprehend the meaning of a word from…

Computation and Language · Computer Science 2017-08-17 Tzu-Ray Su , Hung-Yi Lee

One of the ultimate goals for linguists is to find universal properties in human languages. Although words are generally considered as representing arbitrary mapping between linguistic forms and meanings, we propose a new universal law that…

Computation and Language · Computer Science 2020-05-06 Li-Min Wang , Sun-Ting Tsai , Shan-Jyun Wu , Meng-Xue Tsai , Daw-Wei Wang , Yi-Ching Su , Tzay-Ming Hong

Large language models (LLMs) such as ChatGPT have exhibited remarkable performance in generating human-like texts. However, machine-generated texts (MGTs) may carry critical risks, such as plagiarism issues, misleading information, or…

Computation and Language · Computer Science 2024-03-01 Shuhai Zhang , Yiliao Song , Jiahao Yang , Yuanqing Li , Bo Han , Mingkui Tan

Given a knowledge base KB containing first-order and statistical facts, we consider a principled method, called the random-worlds method, for computing a degree of belief that some formula Phi holds given KB. If we are reasoning about a…

Artificial Intelligence · Computer Science 2009-09-25 A. J. Grove , J. Y. Halpern , D. Koller

We consider the maximum entropy problems associated with R\'enyi $Q$-entropy, subject to two kinds of constraints on expected values. The constraints considered are a constraint on the standard expectation, and a constraint on the…

Information Theory · Computer Science 2008-12-18 Jean-François Bercher

This paper introduces a simple Markov process inspired by the problem of quasicrystal growth. It acts over two-letter words by randomly performing \emph{flips}, a local transformation which exchanges two consecutive different letters. More…

Probability · Mathematics 2010-10-07 Olivier Bodini , Thomas Fernique , Damien Regnault

We investigate the large deviations of the shape of the random RSK Young diagrams associated with a random word of size $n$ whose letters are independently drawn from an alphabet of size $m=m(n)$. When the letters are drawn uniformly and…

Probability · Mathematics 2016-08-14 Christian Houdré , Jinyong Ma

This paper describes the probabilistic behaviour of a random Sturmian word. It performs the probabilistic analysis of the recurrence function which can be viewed as a waiting time to discover all the factors of length $n$ of the Sturmian…

Discrete Mathematics · Computer Science 2016-10-06 Pablo Rotondo , Brigitte Vallee

Kernel methods represent one of the most powerful tools in machine learning to tackle problems expressed in terms of function values and derivatives due to their capability to represent and model complex relations. While these methods show…

Statistics Theory · Mathematics 2015-11-06 Bharath K. Sriperumbudur , Zoltan Szabo

We train a 318M-parameter Transformer language model from scratch on a curated corpus of 1.56 billion tokens of pure Classical Chinese, with zero English characters or Arabic numerals. Through systematic out-of-distribution (OOD) testing,…

Computation and Language · Computer Science 2026-04-17 Jiuting Chen , Yuan Lian , Hao Wu , Tianqi Huang , Hiroshi Sasaki , Makoto Kouno , Jongil Choi

The main theme of this paper is the enumeration of the occurrence of a pattern in words and permutations. We mainly focus on asymptotic properties of the sequence $f_r^v(k,n),$ the number of $n$-array $k$-ary words that contain a given…

Combinatorics · Mathematics 2019-05-15 Toufik Mansour , Reza Rastegar , Alexander Roitershtein