English
Related papers

Related papers: The 'Letter' Distribution in the Chinese Language

200 papers

The statistical methods derived and described in this thesis provide new ways to elucidate the structural properties of text and other symbolic sequences. Generically, these methods allow detection of a difference in the frequency of a…

Computation and Language · Computer Science 2012-07-10 Ted Dunning

The data on 13 typologically different languages have been processed using a two-parameter word length model, based on 1-displaced uniform Poisson distribution. Statistical dependencies of the 2nd parameter on the 1st one are revealed for…

Computation and Language · Computer Science 2007-05-23 Victor Kromer

As the structure of Chinese characters are very different, it is very difficult to input Chinese characters into computer quickly and conveniently. The conventional keyboard does not support the pictorial characters in Chinese language.…

Human-Computer Interaction · Computer Science 2013-10-14 Umakant Mishra

Construction grammar posits that constructions, or form-meaning pairings, are acquired through experience with language (the distributional learning hypothesis). But how much information about constructions does this distribution actually…

Computation and Language · Computer Science 2025-09-26 Joshua Rozner , Leonie Weissweiler , Kyle Mahowald , Cory Shain

Causal processes can give rise to distinctive distributions in the linguistic variables that they affect. Consequently, a secure understanding of a variable's distribution can hold a key to understanding the forces that have causally shaped…

Computation and Language · Computer Science 2021-01-01 Jayden L. Macklin-Cordes , Erich R. Round

Variation in language use, shaped by speakers' sociocultural background and specific context of use, offers a rich lens into cultural perspectives, values, and opinions. For example, Chinese students discuss "healthy eating" with words like…

Computation and Language · Computer Science 2026-04-10 Eylon Caplan , Tania Chakraborty , Dan Goldwasser

In this paper, we compute and demonstrate the equivalence of the joint distribution of the first letter and descent statistics on six avoidance classes of permutations corresponding to two patterns of length four. This distribution is in…

Combinatorics · Mathematics 2021-05-19 Toufik Mansour , Mark Shattuck

We use an information-theoretic measure of linguistic similarity to investigate the organization and evolution of scientific fields. An analysis of almost 20M papers from the past three decades reveals that the linguistic similarity is…

Digital Libraries · Computer Science 2018-01-30 Laercio Dias , Martin Gerlach , Joachim Scharloth , Eduardo G. Altmann

It is intuitive that NLP tasks for logographic languages like Chinese should benefit from the use of the glyph information in those languages. However, due to the lack of rich pictographic evidence in glyphs and the weak generalization…

Computation and Language · Computer Science 2020-05-22 Yuxian Meng , Wei Wu , Fei Wang , Xiaoya Li , Ping Nie , Fan Yin , Muyu Li , Qinghong Han , Xiaofei Sun , Jiwei Li

One important aspect of the relationship between spoken and written Chinese is the ranked syllable-to-character mapping spectrum, which is the ranked list of syllables by the number of characters that map to the syllable. Previously, this…

Computation and Language · Computer Science 2013-08-29 Wentian Li

Linguistics holds unique characteristics of generality, stability, and nationality, which will affect the formulation of extraction strategies and should be incorporated into the relation extraction. Chinese open relation extraction is not…

Computation and Language · Computer Science 2020-10-30 Shengbin Jia

An important body of quantitative linguistics is constituted by a series of statistical laws about language usage. Despite the importance of these linguistic laws, some of them are poorly formulated, and, more importantly, there is no…

Physics and Society · Physics 2020-11-09 Alvaro Corral , Isabel Serra

We show that the predictability of letters in written English texts depends strongly on their position in the word. The first letters are usually the least easy to predict. This agrees with the intuitive notion that words are well defined…

Physics and Society · Physics 2017-04-24 Thomas Schürmann , Peter Grassberger

It was only until the 20th century when the Chinese language began using punctuation. In fact, many ancient Chinese texts contain thousands of lines with no distinct punctuation marks or delimiters in sight. The lack of punctuation in such…

Computation and Language · Computer Science 2024-09-18 Tracy Cai , Kimmy Chang , Fahad Nabi

Active languages such as Bangla (or Bengali) evolve over time due to a variety of social, cultural, economic, and political issues. In this paper, we analyze the change in the written form of the modern phase of Bangla quantitatively in…

Computation and Language · Computer Science 2013-10-08 Paheli Bhattacharya , Arnab Bhattacharya

Zipf's law is a fundamental paradigm in the statistics of written and spoken natural language as well as in other communication systems. We raise the question of the elementary units for which Zipf's law should hold in the most natural way,…

Physics and Society · Physics 2015-07-14 Alvaro Corral , Gemma Boleda , Ramon Ferrer-i-Cancho

Proverbs are an essential component of language and culture, and though much attention has been paid to their history and currency, there has been comparatively little quantitative work on changes in the frequency with which they are used…

Computation and Language · Computer Science 2021-07-13 E. Davis , C. M. Danforth , W. Mieder , P. S. Dodds

We systematically investigate geometric patterns in Chinese character embeddings using PHATE manifold analysis. Through cross-validation across seven embedding models and eight dimensionality reduction methods, we observe clustering…

Computation and Language · Computer Science 2025-10-03 Wen G. Gong

Given the advantage and recent success of English character-level and subword-unit models in several NLP tasks, we consider the equivalent modeling problem for Chinese. Chinese script is logographic and many Chinese logograms are composed…

Computation and Language · Computer Science 2018-09-11 Falcon Z. Dai , Zheng Cai

One of the ultimate goals for linguists is to find universal properties in human languages. Although words are generally considered as representing arbitrary mapping between linguistic forms and meanings, we propose a new universal law that…

Computation and Language · Computer Science 2020-05-06 Li-Min Wang , Sun-Ting Tsai , Shan-Jyun Wu , Meng-Xue Tsai , Daw-Wei Wang , Yi-Ching Su , Tzay-Ming Hong
‹ Prev 1 4 5 6 7 8 10 Next ›