中文
相关论文

相关论文: Revisiting Waiting Times in DNA evolution

200 篇论文

Minimizer schemes, or just minimizers, are a very important computational primitive in sampling and sketching biological strings. Assuming a fixed alphabet of size $\sigma$, a minimizer is defined by two integers $k,w\ge2$ and a total order…

组合数学 · 数学 2024-11-27 Shay Golan , Arseny M. Shur

Every k entries in a permutation can have one of k! different relative orders, called patterns. How many times does each pattern occur in a large random permutation of size n? The distribution of this k!-dimensional vector of pattern…

组合数学 · 数学 2023-09-14 Chaim Even-Zohar

Several important biological processes are initiated by the binding of a protein to a specific site on the DNA. The strategy adopted by a protein, called transcription factor (TF), for searching its specific binding site on the DNA has been…

生物物理 · 物理学 2018-10-17 Soumendu Ghosh , Bhavya Mishra , Anatoly B. Kolomeisky , Debashish Chowdhury

Recent studies reveal even the smallest genomes such as viruses evolve through complex and stochastic processes, and the assumption of independent alleles is not valid in most applications. Advances in sequencing technologies produce…

种群与进化 · 定量生物学 2017-10-30 Hyunjin Shim

Word matches are often used in sequence comparison methods, either as a measure of sequence similarity or in the first search steps of algorithms such as BLAST or BLAT. The D2 statistic is the number of matches of words of k letters between…

定量方法 · 定量生物学 2009-09-09 Sylvain Foret , Susan R. Wilson , Conrad J. Burden

Motivation: The discovery of relationships between gene expression measurements and phenotypic responses is hampered by both computational and statistical impediments. Conventional statistical methods are less than ideal because they either…

统计方法学 · 统计学 2019-07-16 Lei Ding , Daniel J. McDonald

Genomics is changing our understanding of humans, evolution, diseases, and medicines to name but a few. As sequencing technology is developed collecting DNA sequences takes less time thereby generating more genetic data every day. Today the…

定量方法 · 定量生物学 2020-07-29 Sahand Salamat , Tajana Rosing

We study the complexity of determining a winning committee under the Chamberlin--Courant voting rule when voters' preferences are single-crossing on a line, or, more generally, on a median graph (this class of graphs includes, e.g., trees…

计算机科学与博弈论 · 计算机科学 2020-10-20 Andrei Constantinescu , Edith Elkind

The time variation of the rank $k$ of words for six Indo-European languages is obtained using data from Google Books. For low ranks the distinct languages behave differently, maybe due to syntaxis rules, whereas for $k>50$ the law of large…

High read depth can be used to assemble short sequence repeats. The existing genome assemblers fail in repetitive regions of longer than average read. I propose a new algorithm for a DNA assembly which uses the relative frequency of reads…

基因组学 · 定量生物学 2015-01-08 Robert M. Nowak

Motivation: A Genomic Dictionary, i.e., the set of the k-mers appearing in a genome, is a fundamental source of genomic information: its collection is the first step in strategic computational methods ranging from assembly to sequence…

数据结构与算法 · 计算机科学 2022-12-07 Raffaele Giancarlo , Gennaro Grimaudo

Transcription is the first step of gene expression, in which a particular segment of DNA is copied to RNA by the enzyme RNA polymerase (RNAP). Despite many details of the complex interactions between DNA and RNA synthesis disclosed…

生物物理 · 物理学 2020-02-17 Xining Xu , Yunxin Zhang

The prevalent technique for DNA sequencing consists of two main steps: shotgun sequencing, where many randomly located fragments, called reads, are extracted from the overall sequence, followed by an assembly algorithm that aims to…

基因组学 · 定量生物学 2016-01-28 Shirshendu Ganguly , Elchanan Mossel , Miklos Z. Racz

Identifying the number of factors in a high-dimensional factor model has attracted much attention in recent years and a general solution to the problem is still lacking. A promising ratio estimator based on the singular values of the lagged…

统计方法学 · 统计学 2018-01-23 Zeng Li , Qinwen Wang , Jianfeng Yao

Distances between sequences based on their $k$-mer frequency counts can be used to reconstruct phylogenies without first computing a sequence alignment. Past work has shown that effective use of k-mer methods depends on 1) model-based…

种群与进化 · 定量生物学 2017-05-22 Chris Durden , Seth Sullivant

We consider the distributions of the lengths of the longest weakly increasing and strongly decreasing subsequences in words of length N from an alphabet of k letters. We find Toeplitz determinant representations for the exponential…

组合数学 · 数学 2009-07-11 Craig A. Tracy , Harold Widom

We propose a frame-based representation of k-mers for detecting sequencing errors and rare variants in next generation sequencing data obtained from populations of closely related genomes. Frames are sets of non-orthogonal basis functions,…

基因组学 · 定量生物学 2016-04-19 Raunaq Malhotra , Manjari Mukhopadhyay , Mary Poss , Raj Acharya

We study word reconstruction problems. Improving a previous result by P. Fleischmann, M. Lejeune, F. Manea, D. Nowotka and M. Rigo, we prove that, for any unknown word $w$ of length $n$ over an alphabet of cardinality $k$, $w$ can be…

离散数学 · 计算机科学 2023-01-05 Gwenaël Richomme , Matthieu Rosenfeld

Transcription is a complex phenomenon that permits the conversion of genetic information into phenotype by means of an enzyme called RNA polymerase, which erratically moves along and scans the DNA template. We perform Bayesian inference…

生物物理 · 物理学 2023-07-18 Massimo Cavallaro , Yuexuan Wang , Daniel Hebenstreit , Ritabrata Dutta

This report aims to characterise certain sojourn time distributions that naturally arise from semi-Markov models. To this end, it describes a family of discrete distributions that extend the geometric distribution for both finite and…

统计方法学 · 统计学 2022-06-23 Kelli Francis-Staite , Langford White