English
Related papers

Related papers: Comparing reverse complementary genomic words base…

200 papers

Probabilistic context-free grammars (PCFGs) are used to define distributions over strings, and are powerful modelling tools in a number of areas, including natural language processing, software engineering, model checking, bio-informatics,…

Formal Languages and Automata Theory · Computer Science 2014-07-08 Colin de la Higuera , James Scicluna , Mark-Jan Nederhof

This study focuses on an alignment-free sequence comparison method: the number of words of length k shared between two sequences, also known as the D_2 statistic. The advantages of the use of this statistic over alignment-based methods are…

Quantitative Methods · Quantitative Biology 2009-09-08 Sylvain Foret , Susan R. Wilson , Conrad J. Burden

Background: Zipf's discovery that word frequency distributions obey a power law established parallels between biological and physical processes, and language, laying the groundwork for a complex systems perspective on human communication.…

Computation and Language · Computer Science 2009-11-11 Eduardo G. Altmann , Janet B. Pierrehumbert , Adilson E. Motter

Understanding epistasis (genetic interaction) may shed some light on the genomic basis of common diseases, including disorders of maximum interest due to their high socioeconomic burden, like schizophrenia. Distance correlation is an…

Statistical analysis of bacteria genomes texts has been performed on the basis of 20 complete genomes origin from Genebank. It has been revealed that the word ranked distributions are quite well approximated by logarithmic law. Results…

Condensed Matter · Physics 2009-10-31 Olga V. Kirillova

A new family of compound Poisson distribution functions from statistical linguistic is used to study the n-tuples and nucleotide composition features of DNA sequences. The relative frequency distribution of the 6-tuples and 7- tuples…

Statistical Mechanics · Physics 2007-05-23 K. L. Ng , S. P. Li

The distribution of bases spacing in human genome was investigated. An analysis of the frequency of occurrence in the human genome of different sequence lengths flanked by one type of nucleotide was carried out showing that the distribution…

Other Quantitative Biology · Quantitative Biology 2019-04-18 Andrzej Z. Górski , Monika Piwowar

When estimating a phylogeny from a multiple sequence alignment, researchers often assume the absence of recombination. However, if recombination is present, then tree estimation and all downstream analyses will be impacted, because…

Discrepancy measures between probability distributions are at the core of statistical inference and machine learning. In many applications, distributions of interest are supported on different spaces, and yet a meaningful correspondence…

Machine Learning · Computer Science 2021-11-23 Zhengxin Zhang , Youssef Mroueh , Ziv Goldfeld , Bharath K. Sriperumbudur

Long-range correlations are found in symbolic sequences from human language, music and DNA. Determining the span of correlations in dolphin whistle sequences is crucial for shedding light on their communicative complexity. Dolphin whistles…

Neurons and Cognition · Quantitative Biology 2014-12-03 Ramon Ferrer-i-Cancho , Brenda McCowan

The analysis of strings of $n$ random variables with geometric distribution has recently attracted renewed interest: Archibald et al. consider the number of distinct adjacent pairs in geometrically distributed words. They obtain the…

Probability · Mathematics 2024-02-14 Guy Louchard , Werner Schachinger , Mark Daniel Ward

We consider a new family of codes, termed asymmetric Lee distance codes, that arise in the design and implementation of DNA-based storage systems and systems with parallel string transmission protocols. The codewords are defined over a…

Information Theory · Computer Science 2016-12-16 Ryan Gabrys , Han Mao Kiah , Olgica Milenkovic

Frameshift mutations in protein-coding DNA sequences produce a drastic change in the resulting protein sequence, which prevents classic protein alignment methods from revealing the proteins' common origin. Moreover, when a large number of…

Quantitative Methods · Quantitative Biology 2011-01-18 Marta L. Gîrdea , Laurent Noé , Gregory Kucherov

Distance correlation has become an increasingly popular tool for detecting the nonlinear dependence between a pair of potentially high-dimensional random vectors. Most existing works have explored its asymptotic distributions under the null…

Statistics Theory · Mathematics 2021-10-06 Lan Gao , Yingying Fan , Jinchi Lv , Qi-Man Shao

Genome rearrangement has been an active area of research in computational comparative genomics for the last three decades. While initially mostly an interesting algorithmic endeavor, now the practical application by applying rearrangement…

Computational Complexity · Computer Science 2025-07-23 Luís Cunha , Thiago Lopes , Uéverton Souza , Leonard Bohnenkämper , Marília D. V. Braga , Jens Stoye

Chargaff's second parity rule (CSPR) asserts that the frequencies of short polynucleotide chains are the same as those of the complementary reversed chains. Up to now, this hypothesis has only been observed empirically and there is…

Probability · Mathematics 2012-01-04 Andrew Hart , Servet Martínez , Felipe Olmos

Herein it is shown that in order to study the statistical properties of DNA sequences in bacterial chromosomes it suffices to consider only one half of the chromosome because they are similar to its corresponding complementary sequence in…

Genomics · Quantitative Biology 2009-11-10 Marco V. Jose , Tzipe Govezensky , Juan R. Bobadilla

The Universal Similarity Metric (USM) has been demonstrated to give practically useful measures of "similarity" between sequence data. Here we have used the USM as an alternative distance metric in a K-Nearest Neighbours (K-NN) learner to…

Machine Learning · Computer Science 2024-05-13 David Lindsay , Sian Lindsay

Recent research has shown that static word embeddings can encode word frequency information. However, little has been studied about this phenomenon and its effects on downstream tasks. In the present work, we systematically study the…

Computation and Language · Computer Science 2023-10-23 Francisco Valentini , Juan Cruz Sosa , Diego Fernandez Slezak , Edgar Altszyler

Computing the similarity between two probability distributions is a recurring theme across control. We introduce a unified family of distances between the probability distributions of two random variables that is based on the discrepancy…

Systems and Control · Electrical Eng. & Systems 2025-10-03 Alexandros E. Tzikas , Arec Jamgochian , Nazim Kemal Ure , Mykel J. Kochenderfer , Stephen P. Boyd