中文
相关论文

相关论文: Revisiting Waiting Times in DNA evolution

200 篇论文

Described are two algorithms to find long approximate palindromes in a string, for example a DNA sequence. A simple algorithm requires O(n)-space and almost always runs in $O(k.n)$-time where n is the length of the string and k is the…

数据结构与算法 · 计算机科学 2007-05-23 L. Allison

Current popular methods in literature of RNA sequencing normalisation do not account for gene length when compared across samples, whilst adjusting for count biases in the data. This creates a gap in the normalisation as bigger genes in RNA…

其他定量生物学 · 定量生物学 2022-09-02 Hilbert Lam Yuen In , Robbe Pincket

We describe a new tool, KATKA, that stores a phylogenetic tree $T$ such that later, given a pattern $P [1..m]$ and an integer $k$, it can quickly return the root of the smallest subtree of $T$ containing all the genomes in which the $k$-mer…

数据结构与算法 · 计算机科学 2022-08-23 Travis Gagie , Sana Kashgouli , Ben Langmead

In genetic systems there is a non-trivial interface between the sequence of symbols which constitutes the chromosome, or ``genotype'', and the products which this sequence encodes --- the ``phenotype''. This interface can be thought of as a…

adap-org · 物理学 2008-02-03 O. Angeles , C. Stephens , H. Waelbroeck

Polymerization of RNA from a template DNA is carried out by a molecular machine called RNA polymerase (RNAP). It also uses the template as a track on which it moves as a motor utilizing chemical energy input. The time it spends at each…

统计力学 · 物理学 2015-05-13 Tripti Tripathi , Gunter M. Schütz , Debashish Chowdhury

This paper introduces a new family of reconstruction codes which is motivated by applications in DNA data storage and sequencing. In such applications, DNA strands are sequenced by reading some subset of their substrings. While previous…

信息论 · 计算机科学 2022-05-10 Yonatan Yehezkeally , Daniella Bar-Lev , Sagi Marcovich , Eitan Yaakobi

In [GM] Guibert and Mansour studied involutions on n letters avoiding (or containing exactly once) 132 and avoiding (or containing exactly once) an arbitrary pattern on k letters. They also established a bijection between 132-avoiding…

组合数学 · 数学 2007-05-23 O. Guibert , T. Mansour

Literature-based knowledge discovery process identifies the important but implicit relations among information embedded in published literature. Existing techniques from Information Retrieval and Natural Language Processing attempt to…

社会与信息网络 · 计算机科学 2019-11-12 Nazim Choudhury , Fahim Faisal , Matloob Khushi

Gene transcription is a stochastic process that involves thousands of reactions. The first set of these reactions, which happen near a gene promoter, are considered to be the most important in the context of stochastic noise. The most…

分子网络 · 定量生物学 2022-02-01 Jaroslav Albert

DNA sequencing is the basic workhorse of modern day biology and medicine. Shotgun sequencing is the dominant technique used: many randomly located short fragments called reads are extracted from the DNA sequence, and these reads are…

信息论 · 计算机科学 2013-02-15 Abolfazl Motahari , Guy Bresler , David Tse

We study the $k$-Bonacci word over the infinite alphabet $\mathbb{N}$. Since the alphabet is infinite, the usual factor complexity is infinite and does not provide any information. We therefore investigate factor occurrence statistics in…

While most current high-throughput DNA sequencing technologies generate short reads with low error rates, emerging sequencing technologies generate long reads with high error rates. A basic question of interest is the tradeoff between read…

信息论 · 计算机科学 2015-01-27 Ilan Shomorony , Thomas Courtade , David Tse

Reconciling gene trees with a species tree is a fundamental problem to understand the evolution of gene families. Many existing approaches reconcile each gene tree independently. However, it is well-known that the evolution of gene families…

种群与进化 · 定量生物学 2018-06-12 Riccardo Dondi , Manuel Lafond , Celine Scornavacca

String kernels are typically used to compare genome-scale sequences whose length makes alignment impractical, yet their computation is based on data structures that are either space-inefficient, or incur large slowdowns. We show that a…

数据结构与算法 · 计算机科学 2015-02-24 Djamal Belazzougui , Fabio Cunial

A gapped repeat is a factor of the form $uvu$ where $u$ and $v$ are nonempty words. The period of the gapped repeat is defined as $|u|+|v|$. The gapped repeat is maximal if it cannot be extended to the left or to the right by at least one…

形式语言与自动机理论 · 计算机科学 2013-10-01 Roman Kolpakov , Mikhail Podolskiy , Mikhail Posypkin , Nickolay Khrapov

Genome sequencing is the basis for many modern biological and medicinal studies. With recent technological advances, metagenomics has become a problem of interest. This problem entails the analysis and reconstruction of multiple DNA…

概率论 · 数学 2022-01-14 Marlee Herring

Synthesis of DNA molecules offers unprecedented advances in storage technology. Yet, the microscopic world in which these molecules reside induces error patterns that are fundamentally different from their digital counterparts. Hence, to…

信息论 · 计算机科学 2017-08-08 Netanel Raviv , Moshe Schwartz , Eitan Yaakobi

Metagenome, a mixture of different genomes (as a rule, bacterial), represents a pattern, and the analysis of its composition is, currently, one of the challenging problems of bioinformatics. In the present study, the possibility of…

定量方法 · 定量生物学 2016-11-04 Valery Kirzhner , Zeev Volkovich , Renata Avros , Katerina Korenblat

Transcription is a fundamental cellular process, and the first step of gene expression. In human cells, it depends on the binding to chromatin of various proteins, including RNA polymerases and numerous transcription factors (TFs).…

Identifying the genes and mutations that drive the emergence of tumors is a major step to improve understanding of cancer and identify new directions for disease diagnosis and treatment. Despite the large volume of genomics data, the…

机器学习 · 计算机科学 2022-04-05 Renan Andrades , Mariana Recamonde-Mendoza