English
Related papers

Related papers: Comparative statistical analysis of bacteria genom…

200 papers

This paper presents a model-based, unsupervised algorithm for recovering word boundaries in a natural-language text from which they have been deleted. The algorithm is derived from a probability model of the source that generated the text.…

Computation and Language · Computer Science 2007-05-23 Michael R. Brent

In this work we explore the dissimilarity between symmetric word pairs, by comparing the inter-word distance distribution of a word to that of its reversed complement. We propose a new measure of dissimilarity between such distributions.…

We study a minimal model for genome evolution whose elementary processes are single site mutation, duplication and deletion of sequence regions and insertion of random segments. These processes are found to generate long-range correlations…

Genomics · Quantitative Biology 2007-05-23 Philipp W. Messer , Peter F. Arndt , Michael Lässig

This paper develops a theory for characterisation of DNA sequences based on their measure representation. The measures are shown to be random cascades generated by an infinitely divisible distribution. This probability distribution is…

Biological Physics · Physics 2009-11-07 Vo Anh , Ka-Sing Lau , Zu-Guo Yu

In special coordinates (codon position--specific nucleotide frequencies) bacterial genomes form two straight lines in 9-dimensional space: one line for eubacterial genomes, another for archaeal genomes. All the 348 distinct bacterial…

Genomics · Quantitative Biology 2007-11-13 A. N. Gorban , A. Yu. Zinovyev

Bacterial genomes and large-scale computer software projects both consist of a large number of components (genes or software packages) connected via a network of mutual dependencies. Components can be easily added or removed from individual…

Molecular Networks · Quantitative Biology 2013-08-12 Tin Yau Pang , Sergei Maslov

The paper proposes various strategies for sampling text data when performing automatic sentence classification for the purpose of detecting missing bibliographic links. We construct samples based on sentences as semantic units of the text…

Machine Learning · Computer Science 2023-01-05 F. V. Krasnova , I. S. Smaznevicha , E. N. Baskakova

In this article, we investigate the properties of phoneme N-grams across half of the world's languages. We investigate if the sizes of three different N-gram distributions of the world's language families obey a power law. Further, the…

Computation and Language · Computer Science 2014-01-07 Taraka Rama , Lars Borin

Ensembl's human non-coding and protein coding genes are used to automatically find DNA pattern motifs. The Backus-Naur form (BNF) grammar for regular expressions (RE) is used by genetic programming to ensure the generated strings are legal.…

Biomolecules · Quantitative Biology 2010-02-02 W. B. Langdon , Olivia Sanchez Graillet , A. P. Harrison

Finding out statistically significant words in DNA and protein sequences forms the basis for many genetic studies. By applying the maximal entropy principle, we give one systematic way to study the nonrandom occurrence of words in DNA or…

Biological Physics · Physics 2009-11-06 Rui Hu , Bin Wang

We present a theoretical and empirical investigation of the statistical behaviour of the words in a text produced by human language. To this aim, we analyse the word distribution of various texts of Italian language selected from a specific…

Neurons and Cognition · Quantitative Biology 2025-04-15 Diederik Aerts , Jonito Aerts Arguëlles , Lester Beltran , Massimiliano Sassoli de Bianchi , Sandro Sozzo

We investigate the number of inverted repeats observed in 37 complete genomes of bacteria. The number of inverted repeats observed is much higher than expected using Markovian models of DNA sequences in most of the eubacteria. By using the…

Statistical Mechanics · Physics 2007-05-23 Fabrizio Lillo , Salvatore Basile , Rosario N. Mantegna

Statistical properties of the taxonomic classification of human languages are studied. It is shown that, at the highest levels of the taxonomic hierarchy, the frequency of taxon members as a function of the number of languages belonging to…

Adaptation and Self-Organizing Systems · Physics 2007-05-23 Damian H. Zanette

The distributions of the number of occurrences of words (the distributions of words for short) play key roles in information theory, statistics, probability theory, ergodic theory, computer science, and DNA analysis. Bassino et al. 2010 and…

Information Theory · Computer Science 2022-11-16 Hayato Takahashi

We analyze the frequency-rank relationship in sub-vocabularies corresponding to three different grammatical classes (nouns, verbs, and others) in a collection of literary works in English, whose words have been automatically tagged…

Computation and Language · Computer Science 2021-06-11 A. Chacoma , D. H. Zanette

In this paper we introduce a method to detect words or phrases in a given sequence of alphabets without knowing the lexicon. Our linear time unsupervised algorithm relies entirely on statistical relationships among alphabets in the input…

Computation and Language · Computer Science 2013-12-31 Tamal Chowdhury , Rabindra Rakshit , Arko Banerjee

The biological literature is rich with sentences that describe causal relations. Methods that automatically extract such sentences can help biologists to synthesize the literature and even discover latent relations that had not been…

Information Retrieval · Computer Science 2019-04-04 Justin Wood , Nicholas J. Matiasz , Alcino J. Silva , William Hsu , Alexej Abyzov , Wei Wang

In a recent article, Behrens and Vingron (JCB 17, 12, 2010) compute waiting times for k-mers to appear during DNA evolution under the assumption that the considered k-mers do not occur in the initial DNA sequence, an issue arising when…

Discrete Mathematics · Computer Science 2011-12-02 S. Behrens , C. Nicaud , P. Nicodeme

Evolution consists of distinct stages: cosmological, biological, linguistic. Since biology verges on natural sciences and linguistics, we expect that it shares structures and features from both forms of knowledge. Indeed, in DNA we…

Other Quantitative Biology · Quantitative Biology 2019-10-01 Argyris Nicolaidis , Fotis Psomopoulos

A new family of compound Poisson distribution functions from statistical linguistic is used to study the n-tuples and nucleotide composition features of DNA sequences. The relative frequency distribution of the 6-tuples and 7- tuples…

Statistical Mechanics · Physics 2007-05-23 K. L. Ng , S. P. Li