中文
相关论文

相关论文: Revisiting Waiting Times in DNA evolution

200 篇论文

In a recent article, Behrens and Vingron (JCB 17, 12, 2010) compute waiting times for k-mers to appear during DNA evolution under the assumption that the considered k-mers do not occur in the initial DNA sequence, an issue arising when…

离散数学 · 计算机科学 2011-12-02 S. Behrens , C. Nicaud , P. Nicodeme

The frequency distributions of DNA k-mers are shaped by fundamental biological processes and offer a window into genome structure and evolution. Inspired by analogies to natural language, prior studies have attempted to model genomic k-mer…

The utility of DNA sequence substrings (k-mers) in alignment-free phylogenetic classification, including that of bacteria and viruses, is increasingly recognized. However, its biological basis eludes many twenty-first century practitioners.…

种群与进化 · 定量生物学 2019-04-29 Donald R. Forsdyke

One possible explanation for the substantial organismal differences between humans and chimpanzees is that there have been changes in gene regulation. Given what is known about transcription factor binding sites, this motivates the…

概率论 · 数学 2007-05-23 Richard Durrett , Deena Schmidt

The amount of non-unique sequence (non-singletons) in a genome directly affects the difficulty of read alignment to a reference assembly for high throughput-sequencing data. Although a greater length increases the chance for reads being…

基因组学 · 定量生物学 2017-03-03 Wentian Li , Jan Freudenberg , Pedro Miramontes

In this article, we review existing probabilistic models for modeling abundance of fixed-length strings (k-mers) in DNA sequencing data. These models capture dependence of the abundance on various phenomena, such as the size and repeat…

定量方法 · 定量生物学 2022-01-03 Askar Gafurov , Tomáš Vinař , Broňa Brejová

The extraction of $k$-mers is a fundamental component in many complex analyses of large next-generation sequencing datasets, including reads classification in genomics and the characterization of RNA-seq datasets. The extraction of all…

定量方法 · 定量生物学 2021-01-19 Diego Santoro , Leonardo Pellegrina , Fabio Vandin

The analysis of biological sequencing data has been one of the biggest applications of string algorithms. The approaches used in many such applications are based on the analysis of k-mers, which are short fixed-length strings present in a…

数据结构与算法 · 计算机科学 2020-06-15 Rayan Chikhi , Jan Holub , Paul Medvedev

Transcription of the genetic message encoded chemically in the sequence of the DNA template is carried out by a molecular machine called RNA polymerase (RNAP). Backward or forward slippage of the nascent RNA with respect to the DNA template…

生物物理 · 物理学 2017-04-18 Soumendu Ghosh , Shubhadeep Patra , Debashish Chowdhury

Language models, especially transformer-based ones, have achieved colossal success in NLP. To be precise, studies like BERT for NLU and works like GPT-3 for NLG are very important. If we consider DNA sequences as a text written with an…

基因组学 · 定量生物学 2026-01-21 Musa Nuri Ihtiyar , Arzucan Ozgur

Genomic evolution can be viewed as string-editing processes driven by mutations. An understanding of the statistical properties resulting from these mutation processes is of value in a variety of tasks related to biological sequence data,…

信息论 · 计算机科学 2018-12-07 Hao Lou , Farzad Farnoud , Moshe Schwartz , Jehoshua Bruck

One of the ubiquitous representation of long DNA sequence is dividing it into shorter k-mer components. Unfortunately, the straightforward vector encoding of k-mer as a one-hot vector is vulnerable to the curse of dimensionality. Worse yet,…

定量方法 · 定量生物学 2017-01-24 Patrick Ng

A DNA palindrome is a segment of double-stranded DNA sequence with inver- sion symmetry which may form secondary structures conferring significant biolog- ical functions ranging from RNA transcription to DNA replication. To test if the…

应用统计 · 统计学 2011-04-28 I-Ping Tu , Yuan-Fu Huang , Shao-Hsuan Wang

The main theme of this paper is the enumeration of the occurrence of a pattern in words and permutations. We mainly focus on asymptotic properties of the sequence $f_r^v(k,n),$ the number of $n$-array $k$-ary words that contain a given…

组合数学 · 数学 2019-05-15 Toufik Mansour , Reza Rastegar , Alexander Roitershtein

This paper describes the probabilistic behaviour of a random Sturmian word. It performs the probabilistic analysis of the recurrence function which can be viewed as a waiting time to discover all the factors of length $n$ of the Sturmian…

离散数学 · 计算机科学 2016-10-06 Pablo Rotondo , Brigitte Vallee

Various approaches to alignment-free sequence comparison are based on the length of exact or inexact word matches between two input sequences. Haubold {\em et al.} (2009) showed how the average number of substitutions between two DNA…

种群与进化 · 定量生物学 2017-09-06 Burkhard Morgenstern , Svenja Schöbel , Chris-André Leimeister

We propose a lightweight data structure for indexing and querying collections of NGS reads data in main memory. The data structure supports the interface proposed in the pioneering work by Philippe et al. for counting and locating $k$-mers…

数据结构与算法 · 计算机科学 2017-03-03 Tomasz Kowalski , Szymon Grabowski , Sebastian Deorowicz

We show that textual analysis of microbial genomes reveal telling footprints of the early evolution of the genomes. The frequencies of word occurrence of random DNA sequences considered as texts in their four nucleotides are expected to…

生物物理 · 物理学 2007-05-23 Li-Ching Hsieh , Liaofu Luo , HC Lee

In this paper, we develop an explicit formula allowing to compute the first k moments of the random count of a pattern in a multi-states sequence generated by a Markov source. We derive efficient algorithms allowing to deal both with low or…

概率论 · 数学 2012-01-24 Grégory Nuel

In this paper we study the structure of specific linear codes called DNA codes. The first attempts on studying such codes have been proposed over four element rings which are naturally matched with DNA four letters. Later, double (pair) DNA…

组合数学 · 数学 2017-03-31 Fatmanur Gürsoy , Elif Segah Oztas , Irfan Siap
‹ 上一页 1 2 3 10 下一页 ›