中文
相关论文

相关论文: Revisiting Waiting Times in DNA evolution

200 篇论文

We present an efficient algorithm for finding all approximate occurrences of a given pattern $p$ of length $m$ in a text $t$ of length $n$ allowing for translocations of equal length adjacent factors and inversions of factors. The algorithm…

数据结构与算法 · 计算机科学 2013-05-09 Szymon Grabowski , Simone Faro , Emanuele Giaquinta

This paper presents a novel method to segment/decode DNA sequences based on n-grams statistical language model. Firstly, we find the length of most DNA 'words' is 12 to 15 bps by analyzing the genomes of 12 model species. Then we design an…

基因组学 · 定量生物学 2015-03-13 Wang Liang

The most common gene regulation mechanism is when a transcription factor protein binds to a regulatory sequence to increase or decrease RNA transcription. However, transcription factors face two main challenges when searching for these…

定量方法 · 定量生物学 2023-11-21 Lucas Hedström , Ludvig Lizana

Transcription Factors (TFs) are proteins that regulate gene expression. The regulation mechanism is via the binding of a TF to a specific part of the gene associated with it, the TF's target. The target of a specific TF corresponds to a…

统计力学 · 物理学 2022-09-02 Ori Hachmo , Ariel Amir

Heaps' or Herdan's law is a linguistic law describing the relationship between the vocabulary/dictionary size (type) and word counts (token) to be a power-law function. Its existence in genomes with certain definition of DNA words is…

基因组学 · 定量生物学 2024-07-02 Wentian Li , Yannis Almirantis , Astero Provata

Recent studies in DNA sequence classification have leveraged sophisticated machine learning techniques, achieving notable accuracy in categorizing complex genomic data. Among these, methods such as k-mer counting have proven effective in…

基因组学 · 定量生物学 2024-01-26 Şükrü Ozan

Minimizers are sampling schemes with numerous applications in computational biology. Assuming a fixed alphabet of size $\sigma$, a minimizer is defined by two integers $k,w\ge2$ and a linear order $\rho$ on strings of length $k$ (also…

数据结构与算法 · 计算机科学 2025-06-06 Arseny Shur

Next Generation Sequencing (NGS) technologies generate large amounts of short read data for many different organisms. The fact that NGS reads are generally short makes it challenging to assemble the reads and reconstruct the original genome…

基因组学 · 定量生物学 2015-04-07 Jie Ren , Kai Song , Minghua Deng , Gesine Reinert , Charles H. Cannon , Fengzhu Sun

Statistical analysis of bacteria genomes texts has been performed on the basis of 20 complete genomes origin from Genebank. It has been revealed that the word ranked distributions are quite well approximated by logarithmic law. Results…

凝聚态物理 · 物理学 2009-10-31 Olga V. Kirillova

Modern high throughput sequencing technologies like metagenomic sequencing generate millions of sequences which have to be classified based on their taxonomic rank. Modern approaches either apply local alignment and comparison to existing…

基因组学 · 定量生物学 2023-03-14 Wolfgang Fuhl , Susanne Zabel , Kay Nieselt

Strong experimental and theoretical evidence shows that transcription factors and other specific DNA-binding proteins find their sites using a two-mode search: alternating between 3D diffusion through the cell and 1D sliding along the DNA.…

生物大分子 · 定量生物学 2008-06-11 Zeba Wunderlich , Leonid A. Mirny

Many commonly studied species now have more than one chromosome-scale genome assembly, revealing a large amount of genetic diversity previously missed by approaches that map short reads to a single reference. However, many species still…

种群与进化 · 定量生物学 2024-09-19 Miles D. Roberts , Olivia Davis , Emily B. Josephs , Robert J. Williamson

The speed of site-specific binding of transcription factor (TFs) proteins with genomic DNA seems to be strongly retarded by the randomly occurring sequence traps. Traps are those DNA sequences sharing significant similarity with the…

亚细胞过程 · 定量生物学 2016-07-22 G. Niranjani , R. Murugan

For DNA sequences of various species we construct the Google matrix G of Markov transitions between nearby words composed of several letters. The statistical distribution of matrix elements of this matrix is shown to be described by a power…

基因组学 · 定量生物学 2013-05-23 Vivek Kandiah , Dima L. Shepelyansky

We introduce an evolutionary algorithm called recombinator-$k$-means for optimizing the highly non-convex kmeans problem. Its defining feature is that its crossover step involves all the members of the current generation, stochastically…

机器学习 · 计算机科学 2022-02-10 Carlo Baldassi

In response to the evolving landscape of data storage, researchers have increasingly explored non-traditional platforms, with DNA-based storage emerging as a cutting-edge solution. Our work is motivated by the potential of in-vivo DNA…

信息论 · 计算机科学 2024-01-09 Ohad Elishco

This paper addresses the uniform random generation of words from a context-free language (over an alphabet of size $k$), while constraining every letter to a targeted frequency of occurrence. Our approach consists in a multidimensional…

数据结构与算法 · 计算机科学 2010-12-21 Olivier Bodini , Yann Ponty

Feature embedding methods have been proposed in literature to represent sequences as numeric vectors to be used in some bioinformatics investigations, such as family classification and protein structure prediction. Recent theoretical…

String Kernel (SK) techniques, especially those using gapped $k$-mers as features (gk), have obtained great success in classifying sequences like DNA, protein, and text. However, the state-of-the-art gk-SK runs extremely slow when we…

机器学习 · 计算机科学 2017-09-19 Ritambhara Singh , Arshdeep Sekhon , Kamran Kowsari , Jack Lanchantin , Beilun Wang , Yanjun Qi

We present ensemble methods in a machine learning (ML) framework combining predictions from five known motif/binding site exploration algorithms. For a given TF the ensemble starts with position weight matrices (PWM's) for the motif,…

基因组学 · 定量生物学 2018-05-11 Yue Fan , Mark Kon , Charles DeLisi