English
Related papers

Related papers: Comparative statistical analysis of bacteria genom…

200 papers

Counting letters in written texts is a very ancient practice. It has accompanied the development of Cryptology, Quantitative Linguistics, and Statistics. In Cryptology, counting frequencies of the different characters in an encrypted…

History and Overview · Mathematics 2013-12-20 Bernard Ycart

We investigate symbolic sequences and in particular information carriers as e.g. books and DNA--strings. First the higher order Shannon entropies are calculated, a characteristic root law is detected. Then the algorithmic entropy is…

adap-org · Physics 2008-02-03 Werner Ebeling , Alexander Neiman , Thorsten Pöschel

Phylogenetic trees can be reconstructed from the matrix which contains the distances between all pairs of languages in a family. Recently, we proposed a new method which uses normalized Levenshtein distances among words with same meaning…

Computation and Language · Computer Science 2015-05-14 Filippo Petroni , Maurizio Serva

Recent experimental and theoretical approaches have attempted to quantify the physical organization (compaction and geometry) of the bacterial chromosome with its complement of proteins (the nucleoid). The genomic DNA exists in a complex…

DNA is subject to large deformations in a wide range of biological processes. Two key examples illustrate how such deformations influence the readout of the genetic information: the sequestering of eukaryotic genes by nucleosomes, and DNA…

Biomolecules · Quantitative Biology 2015-11-10 Stephanie Johnson , Martin Lindén , Rob Phillips

Based on data from a large-scale experiment with human subjects, we conclude that the logarithm of probability to guess a word in context (unpredictability) depends linearly on the word length. This result holds both for poetry and prose,…

Information Theory · Computer Science 2007-07-16 Dmitrii Manin

We investigate the origin of Zipf's law for words in written texts by means of a stochastic dynamical model for text generation. The model incorporates both features related to the general structure of languages and memory effects inherent…

Statistical Mechanics · Physics 2007-05-23 Damián H. Zanette , Marcelo A. Montemurro

With the development of high throughput sequencing technology, it becomes possible to directly analyze mutation distribution in a genome-wide fashion, dissociating mutation rate measurements from the traditional underlying assumptions.…

Genomics · Quantitative Biology 2015-05-14 D. Parkhomchuk , V. S. Amstislavskiy , A. Soldatov , V. Ogryzko

Periodic reversals of the direction of motion in systems of self-propelled rod shaped bacteria enable them to effectively resolve traffic jams formed during swarming and maximize their swarming rate. In this paper, a connection is found…

Biological Physics · Physics 2015-03-17 Richard Gejji , Pavel M. Lushnikov , Mark Alber

Natural language processing models learn word representations based on the distributional hypothesis, which asserts that word context (e.g., co-occurrence) correlates with meaning. We propose that $n$-grams composed of random character…

Computation and Language · Computer Science 2022-04-21 Mark Chu , Bhargav Srinivasa Desikan , Ethan O. Nadler , D. Ruggiero Lo Sardo , Elise Darragh-Ford , Douglas Guilbeault

This paper studies the effect of linguistic constraints on the large scale organization of language. It describes the properties of linguistic networks built using texts of written language with the words randomized. These properties are…

Computation and Language · Computer Science 2011-02-16 Madhav Krishna , Ahmed Hassan , Yang Liu , Dragomir Radev

Data on the number of Open Reading Frames (ORFs) coded by genomes from the 3 domains of Life show some notable general features including essential differences between the Prokaryotes and Eukaryotes, with the number of ORFs growing linearly…

Genomics · Quantitative Biology 2012-05-31 James L. Friar , Terrance Goldman , Juan Pérez-Mercader

A methodology based upon recurrence quantification analysis is proposed for the study of orthographic structure of written texts. Five different orthographic data sets (20th century Italian poems, 20th century American poems, contemporary…

cmp-lg · Computer Science 2012-08-27 F. Orsucci , K. Walter , A. Giuliani , C. L. Webber, , J. P. Zbilut

Conducting experiments with diverse participants in their native languages can uncover insights into culture, cognition, and language that may not be revealed otherwise. However, conducting these experiments online makes it difficult to…

Computation and Language · Computer Science 2023-02-06 Pol van Rijn , Yue Sun , Harin Lee , Raja Marjieh , Ilia Sucholutsky , Francesca Lanzarini , Elisabeth André , Nori Jacoby

This study presents a fascinating linguistic property related to the number of letters in words and their corresponding numerical values. By selecting any arbitrary word, counting its constituent letters, and subsequently spelling out the…

Computation and Language · Computer Science 2025-03-18 Krishna Chaitanya Polavaram

Finding simple, non-recursive, base noun phrases is an important subtask for many natural language processing applications. While previous empirical methods for base NP identification have been rather complex, this paper instead proposes a…

cmp-lg · Computer Science 2007-05-23 Claire Cardie , David Pierce

String matching algorithm plays the vital role in the Computational Biology. The functional and structural relationship of the biological sequence is determined by similarities on that sequence. For that, the researcher is supposed to aware…

Data Structures and Algorithms · Computer Science 2014-01-30 Pandiselvam. P , Marimuthu. T , Lawrance. R

It is well known that correlations in microarray data represent a serious nuisance deteriorating the performance of gene selection procedures. This paper is intended to demonstrate that the correlation structure of microarray data provides…

Applications · Statistics 2007-12-18 Lev Klebanov , Andrei Yakovlev

For DNA sequences of various species we construct the Google matrix G of Markov transitions between nearby words composed of several letters. The statistical distribution of matrix elements of this matrix is shown to be described by a power…

Genomics · Quantitative Biology 2013-05-23 Vivek Kandiah , Dima L. Shepelyansky

The biological process of gene assembly has been modeled based on three types of string rewriting rules, called string pointer rules, defined on so-called legal strings. It has been shown that reduction graphs, graphs that are based on the…

Logic in Computer Science · Computer Science 2008-11-24 Robert Brijder , Hendrik Jan Hoogeboom