English
Related papers

Related papers: Comparative statistical analysis of bacteria genom…

200 papers

The field of computational linguistics constantly presents new challenges and topics for research. Whether it be analyzing word usage changes over time or identifying relationships between pairs of seemingly unrelated words. To this point,…

Computation and Language · Computer Science 2023-10-27 Saptarshi Sengupta

Statistical studies of languages have focused on the rank-frequency distribution of words. Instead, we introduce here a measure of how word ranks change in time and call this distribution \emph{rank diversity}. We calculate this diversity…

Computation and Language · Computer Science 2015-05-15 Germinal Cocho , Jorge Flores , Carlos Gershenson , Carlos Pineda , Sergio Sánchez

We propose that the distribution of DNA words in genomic sequences can be primarily characterized by a double Pareto-lognormal distribution, which explains lognormal and power-law features found across all known genomes. Such a distribution…

Genomics · Quantitative Biology 2007-05-23 Miklós Csűrös , Laurent Noé , Gregory Kucherov

Word matches are often used in sequence comparison methods, either as a measure of sequence similarity or in the first search steps of algorithms such as BLAST or BLAT. The D2 statistic is the number of matches of words of k letters between…

Quantitative Methods · Quantitative Biology 2009-09-09 Sylvain Foret , Susan R. Wilson , Conrad J. Burden

Sequence comparison is a prerequisite to virtually all comparative genomic analyses. It is often realised by sequence alignment techniques, which are computationally expensive. This has led to increased research into alignment-free…

Data Structures and Algorithms · Computer Science 2018-06-08 Panagiotis Charalampopoulos , Maxime Crochemore , Gabriele Fici , Robert Mercas , Solon P. Pissis

This study focuses on an alignment-free sequence comparison method: the number of words of length k shared between two sequences, also known as the D_2 statistic. The advantages of the use of this statistic over alignment-based methods are…

Quantitative Methods · Quantitative Biology 2009-09-08 Sylvain Foret , Susan R. Wilson , Conrad J. Burden

The distribution of bases spacing in human genome was investigated. An analysis of the frequency of occurrence in the human genome of different sequence lengths flanked by one type of nucleotide was carried out showing that the distribution…

Other Quantitative Biology · Quantitative Biology 2019-04-18 Andrzej Z. Górski , Monika Piwowar

The genetic selection of keywords set, the text frequencies of which are considered as attributes in text classification analysis, has been analyzed. The genetic optimization was performed on a set of words, which is the fraction of the…

Information Retrieval · Computer Science 2012-11-15 Bohdan Pavlyshenko

The phenotype of any organism on earth is, in large part, the consequence of interplay between numerous gene products encoded in the genome, and such interplay between gene products affects the evolutionary fate of the genome itself through…

Molecular Networks · Quantitative Biology 2012-01-04 Pan-Jun Kim , Nathan D. Price

Words are fundamental linguistic units that connect thoughts and things through meaning. However, words do not appear independently in a text sequence. The existence of syntactic rules induces correlations among neighboring words. Using an…

Computation and Language · Computer Science 2023-03-15 David Sanchez , Luciano Zunino , Juan De Gregorio , Raul Toral , Claudio Mirasso

The statistical methods derived and described in this thesis provide new ways to elucidate the structural properties of text and other symbolic sequences. Generically, these methods allow detection of a difference in the frequency of a…

Computation and Language · Computer Science 2012-07-10 Ted Dunning

This work proposes a markovian memoryless model for the DNA that simplifies enormously the complexity of it. We encode nucleotide sequences into symbolic sequences, called words, from which we establish meaningful length of words and group…

Biological Physics · Physics 2015-10-09 Shambhavi Srivastava , Murilo S. Baptista

The genetic code is considered to be universal. In order to test if some statistical properties of the coding bacterial genome were due to inherent properties of the genetic code, we compared the autocorrelation function, the scaling…

Genomics · Quantitative Biology 2009-11-10 Jose A Garcia , Samantha Alvarez , Alejandro Flores , Tzipe Govezensky , Juan R. Bobadilla , Marco V. Jose

The recent dramatic increase in online data availability has allowed researchers to explore human culture with unprecedented detail, such as the growth and diversification of language. In particular, it provides statistical tools to explore…

The rank ordered distribution of the codon usage frequencies for 123 bacteriae is best fitted by a three parameters function that is the sum of a constant, an exponential and a linear term in the rank n. The parameters depend (two…

Genomics · Quantitative Biology 2009-11-11 Luc Frappat , Antonino Sciarrino

As is the case of many signals produced by complex systems, language presents a statistical structure that is balanced between order and disorder. Here we review and extend recent results from quantitative characterisations of the degree of…

Computation and Language · Computer Science 2015-03-05 Marcelo A Montemurro , Damián H Zanette

Shaped by natural selection and other evolutionary forces, an organism's evolutionary history is reflected through its genome sequence, content of functional elements and organization. Consequently, organisms connected through phylogeny,…

Genomics · Quantitative Biology 2023-06-16 Serena Lam , Giorgio Gonnella

In a recent Physical Review Letter, Mantegna et. al., report that certain statistical signatures of natural language can be found in non-coding DNA sequences. In this comment we show that random noise with power-law correlation similar to…

adap-org · Physics 2008-02-03 N. E. Israeloff , M. Kagalenko , K. Chan

Our results demonstrated that a previously reported protein name co-occurrence method (5-mention PubGene) which was not based on a hypothesis testing framework, it is generally statistically more significant than the 99th percentile of…

Digital Libraries · Computer Science 2009-01-05 Maurice HT Ling , Christophe Lefevre , Kevin R. Nicholas

Genes are not located randomly along genomes. Synteny, the conservation of their relative positions in genomes of different species, reflects fundamental constraints on natural evolution. We present approaches to infer pairs of co-localized…

Genomics · Quantitative Biology 2013-07-17 Ivan Junier , Olivier Rivoire