English
Related papers

Related papers: Maximum Entropy, Word-Frequency, Chinese Character…

200 papers

Consider a random word $X^n=(X_1,\ldots ,X_n)$ in an alphabet consisting of $4$ letters, with the letters viewed either as $A$, $U$, $G$ and $C$ (i.e., nucleotides in an RNA sequence) or $\alpha$, $\bar{\alpha}$, $\beta$ and $\bar{\beta}$…

Group Theory · Mathematics 2022-01-20 Siddhartha Gadgil , Manjunath Krishnapur

Understanding the complexity of human language requires an appropriate analysis of the statistical distribution of words in texts. We consider the information retrieval problem of detecting and ranking the relevant words of a text by means…

Computation and Language · Computer Science 2008-06-07 Juan P. Herrera , Pedro A. Pury

Until recently, Chinese texts could not be studied using co-word analysis because the words are not separated by spaces in Chinese (and Japanese). A word can be composed of one or more characters. The online availability of programs that…

Computation and Language · Computer Science 2009-11-10 Loet Leydesdorff , Ping Zhou

The net frequency (NF) of a string, of length $m$, in a text, of length $n$, is the number of occurrences of the string in the text with unique left and right extensions. Recently, Guo et al. [CPM 2024] showed that NF is combinatorially…

Data Structures and Algorithms · Computer Science 2024-08-02 Peaker Guo , Seeun William Umboh , Anthony Wirth , Justin Zobel

Given a random word of size $n$ whose letters are drawn independently from an ordered alphabet of size $m$, the fluctuations of the shape of the random RSK Young tableaux are investigated, when $n$ and $m$ converge together to infinity. If…

Probability · Mathematics 2021-06-08 Jean-Christophe Breton , Christian Houdré

Maximum entropy approach to classification is very well studied in applied statistics and machine learning and almost all the methods that exists in literature are discriminative in nature. In this paper, we introduce a maximum entropy…

Information Theory · Computer Science 2013-12-31 Ambedkar Dukkipati , Gaurav Pandey , Debarghya Ghoshdastidar , Paramita Koley , D. M. V. Satya Sriram

The genetic selection of keywords set, the text frequencies of which are considered as attributes in text classification analysis, has been analyzed. The genetic optimization was performed on a set of words, which is the fraction of the…

Information Retrieval · Computer Science 2012-11-15 Bohdan Pavlyshenko

We investigate symbolic sequences and in particular information carriers as e.g. books and DNA--strings. First the higher order Shannon entropies are calculated, a characteristic root law is detected. Then the algorithmic entropy is…

adap-org · Physics 2008-02-03 Werner Ebeling , Alexander Neiman , Thorsten Pöschel

The thermodynamic maximum principle for the Boltzmann-Gibbs-Shannon (BGS) entropy is reconsidered by combining elements from group and measure theory. Our analysis starts by noting that the BGS entropy is a special case of relative entropy.…

Statistical Mechanics · Physics 2008-11-26 Jörn Dunkel , Peter Talkner , Peter Hänggi

Today's probabilistic language generators fall short when it comes to producing coherent and fluent text despite the fact that the underlying models perform well under standard metrics, e.g., perplexity. This discrepancy has puzzled the…

Computation and Language · Computer Science 2025-06-06 Clara Meister , Tiago Pimentel , Gian Wiher , Ryan Cotterell

The maximum entropy principle advocates to evaluate events' probabilities using a distribution that maximizes entropy among those that satisfy certain expectations' constraints. Such principle can be generalized for arbitrary decision…

Machine Learning · Statistics 2021-12-16 Santiago Mazuelas , Yuan Shen , Aritz Pérez

The problem of compression in standard information theory consists of assigning codes as short as possible to numbers. Here we consider the problem of optimal coding -- under an arbitrary coding scheme -- and show that it predicts Zipf's…

Computation and Language · Computer Science 2020-09-24 Ramon Ferrer-i-Cancho , Christian Bentz , Caio Seguin

The time variation of the rank $k$ of words for six Indo-European languages is obtained using data from Google Books. For low ranks the distinct languages behave differently, maybe due to syntaxis rules, whereas for $k>50$ the law of large…

Physics and Society · Physics 2026-02-04 Germinal Cocho , R. F. Rodríguez , Sergio Sánchez , Jorge Flores , Carlos Pineda , Carlos Gershenson

Distributed representations of words encode lexical semantic information, but what type of information is encoded and how? Focusing on the skip-gram with negative-sampling method, we found that the squared norm of static word embedding…

Computation and Language · Computer Science 2023-11-03 Momose Oyama , Sho Yokoi , Hidetoshi Shimodaira

In this article, how word embeddings can be used as features in Chinese sentiment classification is presented. Firstly, a Chinese opinion corpus is built with a million comments from hotel review websites. Then the word embeddings which…

Computation and Language · Computer Science 2015-11-06 Yiou Lin , Hang Lei , Jia Wu , Xiaoyu Li

We construct the generalized entropy optimized by a given arbitrary statistical distribution with a finite linear expectation value of a random quantity of interest. This offers, via the maximum entropy principle, a unified basis for a…

Statistical Mechanics · Physics 2009-11-07 Sumiyoshi Abe

Automatic character generation is an appealing solution for new typeface design, especially for Chinese typefaces including over 3700 most commonly-used characters. This task has two main pain points: (i) handwritten characters are usually…

Computer Vision and Pattern Recognition · Computer Science 2019-05-07 Chuan Wen , Jie Chang , Ya Zhang , Siheng Chen , Yanfeng Wang , Mei Han , Qi Tian

The Zipf's law is the major regularity of statistical linguistics that served as a prototype for rank-frequency relations and scaling laws in natural sciences. Here we show that the Zipf's law -- together with its applicability for a single…

Data Analysis, Statistics and Probability · Physics 2015-06-15 Armen E. Allahverdyan , Weibing Deng , Q. A. Wang

Let $M(\chi)$ denote the maximum of $|\sum_{n\le N}\chi(n)|$ for a given non-principal Dirichlet character $\chi \pmod q$, and let $N_\chi$ denote a point at which the maximum is attained. In this article we study the distribution of…

Number Theory · Mathematics 2020-06-29 Jonathan Bober , Leo Goldmakher , Andrew Granville , Dimitris Koukoulopoulos

Natural language processing models learn word representations based on the distributional hypothesis, which asserts that word context (e.g., co-occurrence) correlates with meaning. We propose that $n$-grams composed of random character…

Computation and Language · Computer Science 2022-04-21 Mark Chu , Bhargav Srinivasa Desikan , Ethan O. Nadler , D. Ruggiero Lo Sardo , Elise Darragh-Ford , Douglas Guilbeault