English
Related papers

Related papers: Revisiting Waiting Times in DNA evolution

200 papers

A major challenge in next-generation genome sequencing (NGS) is to assemble massive overlapping short reads that are randomly sampled from DNA fragments. To complete assembling, one needs to finish a fundamental task in many leading…

Genomics · Quantitative Biology 2015-05-26 Yang Li , XifengYan

Designing short DNA words is a problem of constructing a set (i.e., code) of n DNA strings (i.e., words) with the minimum length such that the Hamming distance between each pair of words is at least k and the n words satisfy a set of…

Data Structures and Algorithms · Computer Science 2012-02-01 Ming-Yang Kao , Henry C. M. Leung , He Sun , Yong Zhang

Phylogenetically informed k-mers, or phylo-k-mers for short, are k-mers that are predicted to appear within a given genomic region at predefined locations of a fixed phylogeny. Given a reference alignment for this genomic region and…

Quantitative Methods · Quantitative Biology 2022-09-21 Nikolai Romashchenko , Benjamin Linard , Fabio Pardi , Eric Rivals

Let x and y be two length n DNA sequences, and suppose we would like to estimate the divergence time T. A well known simple but crude estimate of T is p := d(x,y)/n, the fraction of mutated sites (the p-distance). We establish a posterior…

Populations and Evolution · Quantitative Biology 2025-07-29 Joseph Mathews , Scott C. Schmidler

The classical pattern matching paradigm is that of seeking occurrences of one string - the pattern, in another - the text, where both strings are drawn from an alphabet set $\Sigma$. Assuming the text length is $n$ and the pattern length is…

Data Structures and Algorithms · Computer Science 2022-07-19 Ora Amir , Amihood Amir , Aviezri Fraenkel , David Sarne

Surviving in a diverse environment requires corresponding organism responses. At the cellular level, such adjustment relies on the transcription factors (TFs) which must rapidly find their target sequences amidst a vast amount of…

Quantitative Methods · Quantitative Biology 2009-11-13 Vahid Rezania , Jack Tuszynski , Michael Hendzel

Unbalanced translocations are among the most frequent chromosomal alterations, accounted for 30\% of all losses of heterozygosity, a major genetic event causing inactivation of tumor suppressor genes. Despite of their central role in…

Data Structures and Algorithms · Computer Science 2018-12-04 Domenico Cantone , Simone Faro , Arianna Pavone

The advent of new experimental genomic technologies and the massive increase of DNA sequence information is helping researchers better understand how our genes work. Recently, experiments on mRNA abundance (gene expression) have revealed…

Biomolecules · Quantitative Biology 2009-11-10 T. Ochiai , J. C. Nacher , T. Akutsu

Although the applications of Non-Homogeneous Poisson Processes to model and study the threshold overshoots of interest in different time series of measurements have proven to provide good results, they needed to be complemented with an…

Applications · Statistics 2023-09-15 Biviana Marcela Suárez-Sierra , Arrigo Coen , Carlos Alberto Taimal

This paper presents a probabilistic approach for DNA sequence analysis. A DNA sequence consists of an arrangement of the four nucleotides A, C, T and G and different representation schemes are presented according to a probability measure…

Quantitative Methods · Quantitative Biology 2010-02-12 Amrita Priyam , B. M. Karan , G. Sahoo

Consider a sequence X_1, X_2,... of i.i.d. uniform random variables taking values in the alphabet set {1,2,...,d}. A k-superpattern is a realization of X_1,...,X_t that contains, as an embedded subsequence, each of the non-order-isomorphic…

Probability · Mathematics 2013-02-20 Anant Godbole , Martha Liendo

For taxonomic classification, we are asked to index the genomes in a phylogenetic tree such that later, given a DNA read, we can quickly choose a small subtree likely to contain the genome from which that read was drawn. Although popular…

Data Structures and Algorithms · Computer Science 2024-04-08 Dominika Draesslerová , Omar Ahmed , Travis Gagie , Jan Holub , Ben Langmead , Giovanni Manzini , Gonzalo Navarro

Tandem duplication in DNA is the process of inserting a copy of a segment of DNA adjacent to the original position. Motivated by applications that store data in living organisms, Jain {\em et al.} (2016) proposed the study of codes that…

Combinatorics · Mathematics 2017-11-20 Yeow Meng Chee , Johan Chrisnata , Han Mao Kiah , Tuan Thanh Nguyen

The distributions of the number of occurrences of words (the distributions of words for short) play key roles in information theory, statistics, probability theory, ergodic theory, computer science, and DNA analysis. Bassino et al. 2010 and…

Information Theory · Computer Science 2022-11-16 Hayato Takahashi

Recent experiments show that transcription factors (TFs) indeed use the facilitated diffusion mechanism to locate their target sequences on DNA in living bacteria cells: TFs alternate between sliding motion along DNA and relocation events…

Quantitative Methods · Quantitative Biology 2015-07-10 Maximilian Bauer , Emil S. Rasmussen , Michael A. Lomholt , Ralf Metzler

K-mer counting is a requisite process for DNA assembly because it speeds up its overall process. The frequency of K-mers is used for estimating the parameters of DNA assembly, error correction, etc. The process also provides a list of…

Databases · Computer Science 2023-05-15 Sabuzima Nayak , Ripon Patgiri

The Moran model with recombination is considered, which describes the evolution of the genetic composition of a population under recombination and resampling. There are $n$ sites (or loci), a finite number of letters (or alleles) at every…

Probability · Mathematics 2018-11-01 Mareike Esser , Sebastian Probst , Ellen Baake

Objections to Darwinian evolution are often based on the time required to carry out the necessary mutations. Seemingly, exponential numbers of mutations are needed. We show that such estimates ignore the effects of natural selection, and…

Probability · Mathematics 2015-05-20 Herbert S. Wilf , Warren J. Ewens

To ensure fast gene activation, Transcription Factors (TF) use a mechanism known as facilitated diffusion to find their DNA promoter site. Here we analyze such a process where a TF alternates between 3D and 1D diffusion. In the latter (TF…

Subcellular Processes · Quantitative Biology 2015-05-27 Juergen Reingruber , David Holcman

We prove that a random word of length $n$ over a $k$-ary fixed alphabet contains, on expectation, $\Theta(\sqrt{n})$ distinct palindromic factors. We study this number of factors, $E(n,k)$, in detail, showing that the limit…

Combinatorics · Mathematics 2016-09-13 Mikhail Rubinchik , Arseny M. Shur