English
Related papers

Related papers: Even better correction of genome sequencing data

200 papers

Variational formulations of reconstruction in computed tomography have the notable drawback of requiring repeated evaluations of both the forward Radon transform and either its adjoint or an approximate inverse transform which are…

Numerical Analysis · Mathematics 2017-05-23 Richard C. Barnard , Rick Archibald

We present a novel scheme to boost detection power for kernel maximum mean discrepancy based sequential change-point detection procedures. Our proposed scheme features an optimal sub-sampling of the history data before the detection…

Methodology · Statistics 2023-01-19 Song Wei , Chaofan Huang

Given natural limitations on the length DNA sequences, designing phylogenetic reconstruction methods which are reliable under limited information is a crucial endeavor. There have been two approaches to this problem: reconstructing partial…

Data Structures and Algorithms · Computer Science 2008-12-10 Radu Mihaescu , Cameron Hill , Satish Rao

The mRNA optimization is critical for therapeutic and biotechnological applications, since sequence features directly govern protein expression levels and efficacy. However, current methods face significant challenges in simultaneously…

Quantitative Methods · Quantitative Biology 2025-06-02 Zheng Gong , Ziyi Jiang , Weihao Gao , Deng Zhuo , Lan Ma

Next generation sequencing technology rapidly produces massive volume of data and quality control of this sequencing data is essential to any genomic analysis. Here we present MEEPTOOLS, which is a collection of open-source tools based on…

Genomics · Quantitative Biology 2015-12-11 Vishal N. Koparde , Hardik I. Parikh , Steven P. Bradley , Nihar U. Sheth

Chinese grammatical error correction (CGEC) faces serious overcorrection challenges when employing autoregressive generative models such as sequence-to-sequence (Seq2Seq) models and decoder-only large language models (LLMs). While previous…

Computation and Language · Computer Science 2024-06-04 Haihui Yang , Xiaojun Quan

Advances in high-throughput sequencing technology have led to significant progress in measuring gene expressions at the single-cell level. The amount of publicly available single-cell RNA-seq (scRNA-seq) data is already surpassing 50M…

Machine Learning · Computer Science 2024-02-27 Jing Gong , Minsheng Hao , Xingyi Cheng , Xin Zeng , Chiming Liu , Jianzhu Ma , Xuegong Zhang , Taifeng Wang , Le Song

We present a novel technique for automatic program correction in MOOCs, capable of fixing both syntactic and semantic errors without manual, problem specific correction strategies. Given an incorrect student program, it generates candidate…

Programming Languages · Computer Science 2016-07-12 Yewen Pu , Karthik Narasimhan , Armando Solar-Lezama , Regina Barzilay

Summary: FermiKit is a variant calling pipeline for Illumina data. It de novo assembles short reads and then maps the assembly against a reference genome to call SNPs, short insertions/deletions (INDELs) and structural variations (SVs).…

Genomics · Quantitative Biology 2015-04-27 Heng Li

New long read sequencing technologies, like PacBio SMRT and Oxford NanoPore, can produce sequencing reads up to 50,000 bp long but with an error rate of at least 15%. Reducing the error rate is necessary for subsequent utilisation of the…

Genomics · Quantitative Biology 2021-11-18 Leena Salmela , Riku Walve , Eric Rivals , Esko Ukkonen

Forecasting methods are affected by data quality issues in two ways: 1. they are hard to predict, and 2. they may affect the model negatively when it is updated with new data. The latter issue is usually addressed by pre-processing the data…

Machine Learning · Computer Science 2024-04-30 Rodrigo Tuna , Yassine Baghoussi , Carlos Soares , João Mendes-Moreira

Current computational methods for exon-intron structure prediction from a cluster of transcript (EST, mRNA) data do not exhibit the time and space efficiency necessary to process large clusters of over than 20,000 ESTs and genes longer than…

Genomics · Quantitative Biology 2010-05-11 Paola Bonizzoni , Gianluca Della Vedova , Yuri Pirola , Raffaella Rizzi

The decreasing costs and increasing speed and accuracy of DNA sample collection, preparation, and sequencing has rapidly produced an enormous volume of genetic data. However, fast and accurate analysis of the samples remains a bottleneck.…

Quantitative Methods · Quantitative Biology 2017-04-13 Stephanie Dodson , Darrell O. Ricke , Jeremy Kepner , Nelson Chiu , Anna Shcherbina

Genome rearrangement distances are an established method in genome comparison. Works in this area may include various rearrangement operations representing large-scale mutations, gene orientation information, the number of nucleotides in…

Data Structures and Algorithms · Computer Science 2026-01-01 Gabriel Siqueira , Alexsandro Oliveira Alexandrino , Zanoni Dias

Traditionally, we usually utilize the method of shotgun to cut a DNA sequence into pieces and we have to reconstruct the original DNA sequence from the pieces, those are widely used method for DNA assembly. Emerging DNA sequence…

Distributed, Parallel, and Cluster Computing · Computer Science 2014-04-15 Yukun Zhong , ZhiWei He , XianHong Wang , XiongBin Cao

While most current high-throughput DNA sequencing technologies generate short reads with low error rates, emerging sequencing technologies generate long reads with high error rates. A basic question of interest is the tradeoff between read…

Information Theory · Computer Science 2015-01-27 Ilan Shomorony , Thomas Courtade , David Tse

Deciphering cell type heterogeneity is crucial for systematically understanding tissue homeostasis and its dysregulation in diseases. Computational deconvolution is an efficient approach estimating cell type abundances from a variety of…

Other Quantitative Biology · Quantitative Biology 2023-09-06 Lana X. Garmire , Yijun Li , Qianhui Huang , Chuan Xu , Sarah Teichmann , Naftali Kaminski , Matteo Pellegrini , Quan Nguyen , Andrew E. Teschendorff

Sequence classification has numerous applications in various fields. Despite extensive studies in the last decades, many challenges still exist, particularly in pattern-based methods. Existing pattern-based methods measure the…

Machine Learning · Computer Science 2023-10-23 Junjie Dong , Mudi Jiang , Lianyu Hu , Zengyou He

In-Memory Computing (IMC) introduces a new paradigm of computation that offers high efficiency in terms of latency and power consumption for AI accelerators. However, the non-idealities and defects of emerging technologies used in advanced…

Efficient and accurate low-rank approximations of multiple data sources are essential in the era of big data. The scaling of kernel-based learning algorithms to large datasets is limited by the O(n^2) computation and storage complexity of…

Machine Learning · Computer Science 2020-12-10 Martin Stražar , Tomaž Curk
‹ Prev 1 8 9 10 Next ›