English
Related papers

Related papers: Rank-frequency relation for Chinese characters

200 papers

Zipf's law is shown to arise as the variational solution of a problem formulated in Fisher's terms. An appropriate minimization process involving Fisher information and scale-invariance yields this universal rank distribution. As an example…

Physics and Society · Physics 2015-05-13 A. Hernando , D. Puigdomenech , D. Villuendas , C. Vesperinas , A. Plastino

The origin of long-range letter correlations in natural texts is studied using random walk analysis and Jensen-Shannon divergence. It is concluded that they result from slow variations in letter frequency distribution, which are a…

Computation and Language · Computer Science 2016-11-27 Dmitrii Y. Manin

In the setting of minimal local grammar-based coding, the input string is represented as a grammar with the minimal output length defined via simple symbol-by-symbol encoding. This paper discusses four contributions to this field. First, we…

Information Theory · Computer Science 2025-04-17 Łukasz Dębowski

The mapping of lexical meanings to wordforms is a major feature of natural languages. While usage pressures might assign short words to frequent meanings (Zipf's law of abbreviation), the need for a productive and open-ended vocabulary,…

Computation and Language · Computer Science 2021-05-04 Tiago Pimentel , Irene Nikkarinen , Kyle Mahowald , Ryan Cotterell , Damián Blasi

We show how generalized Gibbs-Shannon entropies can provide new insights on the statistical properties of texts. The universal distribution of word frequencies (Zipf's law) implies that the generalized entropies, computed at the word level,…

Physics and Society · Physics 2017-02-15 Eduardo G. Altmann , Laercio Dias , Martin Gerlach

The performance of deep learning in natural language processing has been spectacular, but the reasons for this success remain unclear because of the inherent complexity of deep learning. This paper provides empirical evidence of its…

Computation and Language · Computer Science 2018-02-07 Shuntaro Takahashi , Kumiko Tanaka-Ishii

Words in some natural languages can have a composite structure. Elements of this structure include the root (that could also be composite), prefixes and suffixes with which various nuances and relations to other words can be expressed.…

Computation and Language · Computer Science 2017-09-05 Rustem Takhanov , Zhenisbek Assylbekov

This paper mainly discusses the generation of personalized fonts as the problem of image style transfer. The main purpose of this paper is to design a network framework that can extract and recombine the content and style of the characters.…

Computer Vision and Pattern Recognition · Computer Science 2020-04-08 Fenxi Xiao , Jie Zhang , Bo Huang , Xia Wu

The results of quantitative analysis of word distribution in two fables in Ukrainian by Ivan Franko: "Mykyta the Fox" and "Abu-Kasym's slippers" are reported. Our study consists of two parts: the analysis of frequency-rank distributions and…

Data Analysis, Statistics and Probability · Physics 2009-04-03 Yu. Holovatch , V. Palchykov

Zipf's law, originally discovered in natural language and later generalized to the Zipf-Mandelbrot law, describes a power-law relationship between the frequency of a Zipfian element and its rank. Due to the semantic characteristics of this…

Applications · Statistics 2026-02-17 Byeongchan Choi , Junwon You , Myung Ock Kim , Jae-Hun Jung

Words follow the law of brevity, i.e. more frequent words tend to be shorter. From a statistical point of view, this qualitative definition of the law states that word length and word frequency are negatively correlated. Here the recent…

Neurons and Cognition · Quantitative Biology 2014-12-03 Stuart Semple , Minna J. Hsu , Govindasamy Agoramoorthy , Ramon Ferrer-i-Cancho

Synthesizing Chinese characters with consistent style using few stylized examples is challenging. Existing models struggle to generate arbitrary style characters with limited examples. In this paper, we propose the Generalized W-Net, a…

Computer Vision and Pattern Recognition · Computer Science 2024-06-12 Haochuan Jiang , Guanyu Yang , Fei Cheng , Kaizhu Huang

Let $\Sigma$ be a countable alphabet. For $r\geq 1$, an infinite sequence $s$ with characters from $\Sigma$ is called $r$-quasi-regular, if for each $\sigma\in\Sigma$ the ratio of the longest to shortest interval between consecutive…

Combinatorics · Mathematics 2019-10-01 Joshua Frisch , Wade Hann-Caruthers , Pooya Vahidi Ferdowsi

We study the rank distribution, the cumulative probability, and the probability density of returns of stock prices of listed firms traded in four stock markets. We find that the rank distribution and the cumulative probability of stock…

Other Condensed Matter · Physics 2008-12-02 Kyungsik Kim , S. -M. Yoon , K. H. Chang

Short text matching often faces the challenges that there are great word mismatch and expression diversity between the two texts, which would be further aggravated in languages like Chinese where there is no natural space to segment words…

Computation and Language · Computer Science 2019-02-26 Yuxuan Lai , Yansong Feng , Xiaohan Yu , Zheng Wang , Kun Xu , Dongyan Zhao

Heaps' or Herdan's law characterizes the word-type vs. word-token relation by a power-law function, which is concave in linear-linear scale but a straight line in log-log scale. However, it has been observed that even in log-log scale, the…

Computation and Language · Computer Science 2026-05-27 Oscar Fontanelli , Wentian Li

Analogical reasoning is effective in capturing linguistic regularities. This paper proposes an analogical reasoning task on Chinese. After delving into Chinese lexical knowledge, we sketch 68 implicit morphological relations and 28 explicit…

Computation and Language · Computer Science 2018-08-27 Shen Li , Zhe Zhao , Renfen Hu , Wensi Li , Tao Liu , Xiaoyong Du

Recognition of Off-line Chinese characters is still a challenging problem, especially in historical documents, not only in the number of classes extremely large in comparison to contemporary image retrieval methods, but also new unseen…

Computer Vision and Pattern Recognition · Computer Science 2018-08-29 Sheng He , Lambert Schomaker

The time variation of the rank $k$ of words for six Indo-European languages is obtained using data from Google Books. For low ranks the distinct languages behave differently, maybe due to syntaxis rules, whereas for $k>50$ the law of large…

Physics and Society · Physics 2026-02-04 Germinal Cocho , R. F. Rodríguez , Sergio Sánchez , Jorge Flores , Carlos Pineda , Carlos Gershenson

Why does Zipf's law give a good description of data from seemingly completely unrelated phenomena? Here it is argued that the reason is that they can all be described as outcomes of a ubiquitous random group division: the elements can be…

Physics and Society · Physics 2015-03-19 Seung Ki Baek , Sebastian Bernhardsson , Petter Minnhagen
‹ Prev 1 8 9 10 Next ›