中文
相关论文

相关论文: Do Read Errors Matter for Genome Assembly?

200 篇论文

De novo genome assembly is challenging in highly repetitive regions; however, reference-guided assemblers often suffer from bias. We propose a framework for pangenome-guided sequence assembly, which can resolve short-read data in complex…

量子物理 · 物理学 2026-02-11 Josh Cudby , James Bonfield , Chenxi Zhou , Richard Durbin , Sergii Strelchuk

As sequencing technologies become more affordable and genomic databases expand continuously, the reuse of publicly available sequencing data emerges as a powerful strategy for studying microbial pathogens. Indeed, raw sequencing reads…

定量方法 · 定量生物学 2025-05-16 Damien Richard , Nils Poulicard

Genome sequence analysis has enabled significant advancements in medical and scientific areas such as personalized medicine, outbreak tracing, and the understanding of evolution. Unfortunately, it is currently bottlenecked by the…

Nanopore sequencing can read substantially longer sequences of nucleic acid molecules, called reads, than other sequencing methods, which has led to advances in genomic analysis such as the gapless human genome assembly. By analyzing the…

基因组学 · 定量生物学 2026-05-21 Simon Ambrozak , Ulysse McConnell , Bhargav Srinivasan , Burak Ozkan , Ernest Zhang , Can Firtina

The rapid advancement of DNA sequencing has produced vast genomic datasets, yet interpreting and engineering genomic function remain fundamental challenges. Recent large language models have opened new avenues for genomic analysis, but…

Due to its higher data density, longevity, energy efficiency, and ease of generating copies, DNA is considered a promising storage technology for satisfying future needs. However, a diverse set of errors including deletions, insertions,…

信息论 · 计算机科学 2022-08-05 Yuanyuan Tang , Shuche Wang , Hao Lou , Ryan Gabrys , Farzad Farnoud

DNA-based storage is an emerging storage technology that provides high information density and long duration. Due to the physical constraints in the reading and writing processes, error correction in DNA storage poses several interesting…

信息论 · 计算机科学 2023-10-04 Jin Sima , Netanel Raviv , Moshe Schwartz , Jehoshua Bruck

We present a parallel algorithm and scalable implementation for genome analysis, specifically the problem of finding overlaps and alignments for data from "third generation" long read sequencers. While long sequences of DNA offer enormous…

分布式、并行与集群计算 · 计算机科学 2020-01-29 Marquita Ellis , Giulia Guidi , Aydın Buluç , Leonid Oliker , Katherine Yelick

Although generative models have made remarkable progress in recent years, their use in critical applications has been hindered by an inability to reliably evaluate the quality of their generated samples. Quality refers to at least two…

机器学习 · 计算机科学 2026-02-18 Nicolas Salvy , Hugues Talbot , Bertrand Thirion

Current techniques in sequencing a genome allow a service provider (e.g. a sequencing company) to have full access to the genome information, and thus the privacy of individuals regarding their lifetime secret is violated. In this paper, we…

基因组学 · 定量生物学 2018-11-28 Ali Gholami , Mohammad Ali Maddah-Ali , Seyed Abolfazl Motahari

Grammar Error Correction(GEC) mainly relies on the availability of high quality of large amount of synthetic parallel data of grammatically correct and erroneous sentence pairs. The quality of the synthetic data is evaluated on how well the…

计算与语言 · 计算机科学 2022-11-01 Vanya Bannihatti Kumar

DNA is a promising storage medium, but its stability and occurrence of Indel errors pose a significant challenge. The relative occurrence of Guanine(G) and Cytosine(C) in DNA is crucial for its longevity, and reverse complementary base…

信息论 · 计算机科学 2024-01-15 NallappaBhavithran G , Selvakumar R

Long-read sequencing has enabled the de novo assembly of several mammalian genomes, but with high cost in computing. Here, we demonstrated de novo assembly of mammalian genome using long reads in an efficient and inexpensive workstation.

基因组学 · 定量生物学 2017-03-31 Hikoyu Suzuki , Norichika Ogata

We investigate the fundamental limits of the recently proposed random access coverage depth problem for DNA data storage. Under this paradigm, it is assumed that the user information consists of $k$ information strands, which are encoded…

信息论 · 计算机科学 2025-09-25 Anina Gruica , Daniella Bar-Lev , Alberto Ravagnani , Eitan Yaakobi

Gene finding is the task of identifying the locations of coding sequences within the vast amount of genetic code contained in the genome. With an ever increasing quantity of raw genome sequences, gene finding is an important avenue towards…

基因组学 · 定量生物学 2025-05-07 Frederikke I. Marin , Dennis Pultz , Wouter Boomsma

Sequencing technologies are prone to errors, making error correction (EC) necessary for downstream applications. EC tools need to be manually configured for optimal performance. We find that the optimal parameters (e.g., k-mer size) are…

基因组学 · 定量生物学 2021-12-21 Atul Sharma , Pranjal Jain , Ashraf Mahgoub , Zihan Zhou , Kanak Mahadik , Somali Chaterji

Sequencing a genome to determine an individual's DNA produces an enormous number of short nucleotide subsequences known as reads, which must be reassembled to reconstruct the full genome. We present a method for analyzing this type of data…

机器学习 · 计算机科学 2025-05-23 Filip Thor , Carl Nettelblad

The synthesis of DNA strands remains the most costly part of the DNA storage system. Thus, to make DNA storage system more practical, the time and materials used in the synthesis process have to be optimized. We consider the most common…

信息论 · 计算机科学 2023-05-15 Johan Chrisnata , Han Mao Kiah , Van Long Phuoc Pham

High-throughput shotgun sequence data makes it possible in principle to accurately estimate population genetic parameters without confounding by SNP ascertainment bias. One such statistic of interest is the proportion of heterozygous sites…

种群与进化 · 定量生物学 2012-12-18 Katarzyna Bryc , Nick Patterson , David Reich

High throughput sequencing is a technology that allows for the generation of millions of reads of genomic data regarding a study of interest, and data from high throughput sequencing platforms are usually count compositions. Subsequent…

定量方法 · 定量生物学 2017-04-07 Jia R. Wu , Jean M. Macklaim , Briana L. Genge , Gregory B. Gloor