English
Related papers

Related papers: Statistical linguistic study of DNA sequences

200 papers

We show that textual analysis of microbial genomes reveal telling footprints of the early evolution of the genomes. The frequencies of word occurrence of random DNA sequences considered as texts in their four nucleotides are expected to…

Biological Physics · Physics 2007-05-23 Li-Ching Hsieh , Liaofu Luo , HC Lee

This paper presents a probabilistic approach for DNA sequence analysis. A DNA sequence consists of an arrangement of the four nucleotides A, C, T and G and different representation schemes are presented according to a probability measure…

Quantitative Methods · Quantitative Biology 2010-02-12 Amrita Priyam , B. M. Karan , G. Sahoo

Sequencing by synthesis is used in many next-generation DNA sequencing technologies. Some of the technologies, especially those exploring the principle of single-molecule sequencing, allow incomplete nucleotide incorporation in each cycle.…

Genomics · Quantitative Biology 2024-05-28 Yong Kong

A family of consistent tests, derived from a characterization of the probability generating function, is proposed for assessing Poissonity against a wide class of count distributions, which includes some of the most frequently adopted…

Statistics Theory · Mathematics 2024-06-11 Antonio Di Noia , Marzia Marcheselli , Caterina Pisani , Luca Pratelli

Much of the on-going statistical analysis of DNA sequences is focused on the estimation of characteristics of coding and non-coding regions that would possibly allow discrimination of these regions. In the current approach, we concentrate…

Genomics · Quantitative Biology 2009-11-10 D. Kugiumtzis , A. Provata

This paper presents a novel method to segment/decode DNA sequences based on n-grams statistical language model. Firstly, we find the length of most DNA 'words' is 12 to 15 bps by analyzing the genomes of 12 model species. Then we design an…

Genomics · Quantitative Biology 2015-03-13 Wang Liang

In this article, we investigate the properties of phoneme N-grams across half of the world's languages. We investigate if the sizes of three different N-gram distributions of the world's language families obey a power law. Further, the…

Computation and Language · Computer Science 2014-01-07 Taraka Rama , Lars Borin

The statistical methods derived and described in this thesis provide new ways to elucidate the structural properties of text and other symbolic sequences. Generically, these methods allow detection of a difference in the frequency of a…

Computation and Language · Computer Science 2012-07-10 Ted Dunning

The codons, sixtyfour in number, are distributed over the coding parts of DNA sequences. The distribution function is the plot of frequency-versus-rank of the codons. These distributions are characterised by parameters that are almost…

Biological Physics · Physics 2009-11-07 A. Som , S. Chattopadhyay , J. Chakrabarti , D. Bandyopadhyay

In this work we seek clusters of genomic words in human DNA by studying their inter-word lag distributions. Due to the particularly spiked nature of these histograms, a clustering procedure is proposed that first decomposes each…

Applications · Statistics 2021-01-13 Ana Helena Tavares , Jakob Raymaekers , Peter J. Rousseeuw , Paula Brito , Vera Afreixo

A fundamental characteristic of natural language is the high rate at which speakers produce novel expressions. Because of this novelty, a heavy-tail of rare events accounts for a significant amount of the total probability mass of…

Computation and Language · Computer Science 2022-03-25 Benjamin LeBrun , Alessandro Sordoni , Timothy J. O'Donnell

The evolution in coding DNA sequences brings new flexibility and freedom to the codon words, even as the underlying nucleotides get significantly ordered. These curious contra-rules of gene organisation are observed from the distribution of…

Biological Physics · Physics 2007-05-23 Sujay Chattopadhyay , William A. Kanner , Jayprokas Chakrabarti

Background: We study the statistical properties of fragment coverage in genome sequencing experiments. In an extension of the classic Lander-Waterman model, we consider the effect of the length distribution of fragments. We also introduce…

Genomics · Quantitative Biology 2010-05-03 Steven N. Evans , Valerie Hower , Lior Pachter

The distributions of the number of occurrences of words (the distributions of words for short) play key roles in information theory, statistics, probability theory, ergodic theory, computer science, and DNA analysis. Bassino et al. 2010 and…

Information Theory · Computer Science 2022-11-16 Hayato Takahashi

We examine the problem of family size statistics (the number of individuals carrying the same surname, or the same DNA sequence) in a given size subsample of an exponentially growing population. We approach the problem from two directions.…

Populations and Evolution · Quantitative Biology 2009-03-24 Yosef E. Maruvka , Nadav M. Shnerb , David A. Kessler

Recent observations in the theory of verse and empirical metrics have suggested that constructing a verse line involves a pattern-matching search through a source text, and that the number of found elements (complete words totaling a…

cmp-lg · Computer Science 2007-05-23 Hideaki Aoyama , John Constable

We propose that the distribution of DNA words in genomic sequences can be primarily characterized by a double Pareto-lognormal distribution, which explains lognormal and power-law features found across all known genomes. Such a distribution…

Genomics · Quantitative Biology 2007-05-23 Miklós Csűrös , Laurent Noé , Gregory Kucherov

A new generalization of the family of Poisson-G is called beta Poisson-G family of distribution. Useful expansions of the probability density function and the cumulative distribution function of the proposed family are derived and seen as…

Statistics Theory · Mathematics 2020-05-22 Laba Handique , Subrata Chakraborty , Farrukh Jamal

In this article we survey properties of mixed Poisson distributions and probabilistic aspects of the Stirling transform: given a non-negative random variable $X$ with moment sequence $(\mu_s)_{s\in\mathbb{N}}$ we determine a discrete random…

Combinatorics · Mathematics 2014-09-12 Markus Kuba , Alois Panholzer

Using recent results on the occurrence times of a string of symbols in a stochastic process with mixing properties, we present a new method for the search of rare words in biological sequences generally modelled by a Markov chain. We obtain…

Probability · Mathematics 2007-11-16 Nicolas Vergne , Miguel Abadi
‹ Prev 1 2 3 10 Next ›