English
Related papers

Related papers: Misassembly Detection using Paired-End Sequence Re…

200 papers

State-of-the-art causal discovery methods usually assume that the observational data is complete. However, the missing data problem is pervasive in many practical scenarios such as clinical trials, economics, and biology. One…

Machine Learning · Computer Science 2023-01-18 Erdun Gao , Ignavier Ng , Mingming Gong , Li Shen , Wei Huang , Tongliang Liu , Kun Zhang , Howard Bondell

Over the past few years, there has been a significant improvement in the domain of few-shot learning. This learning paradigm has shown promising results for the challenging problem of anomaly detection, where the general task is to deal…

Computer Vision and Pattern Recognition · Computer Science 2022-11-24 Soumyajit Karmakar , Abeer Banerjee , Prashant Sadashiv Gidde , Sumeet Saurav , Sanjay Singh

In this paper, fundamental limits in sequencing of a set of closely related DNA molecules are addressed. This problem is called pooled-DNA sequencing which encompasses many interesting problems such as haplotype phasing, metageomics, and…

Information Theory · Computer Science 2016-04-20 Amir Najafi , Damoun Nashta-ali , Seyed Abolfazl Motahari , Mehrdad Khani , Babak H. Khalaj , Hamid R. Rabiee

Short-read DNA sequencing instruments can yield over 1e+12 bases per run, typically composed of reads 150 bases long. Despite this high throughput, de novo assembly algorithms have difficulty reconstructing contiguous genome sequences using…

Genomics · Quantitative Biology 2023-06-09 Eric Chen , Justin Chu , Jessica Zhang , Rene L. Warren , Inanc Birol

Semi-labeled trees are phylogenies whose internal nodes may be labeled by higher-order taxa. Thus, a leaf labeled Mus musculus could nest within a subtree whose root node is labeled Rodentia, which itself could nest within a subtree whose…

Data Structures and Algorithms · Computer Science 2016-05-09 Yun Deng , David Fernández-Baca

In this paper, we solve the sample shortage problem in the human parsing task. We begin with the self-learning strategy, which generates pseudo-labels for unlabeled data to retrain the model. However, directly using noisy pseudo-labels will…

Computer Vision and Pattern Recognition · Computer Science 2020-04-20 Tao Li , Zhiyuan Liang , Sanyuan Zhao , Jiahao Gong , Jianbing Shen

We develop statistically based methods to detect single nucleotide DNA mutations in next generation sequencing data. Sequencing generates counts of the number of times each base was observed at hundreds of thousands to billions of genome…

Applications · Statistics 2012-10-01 Omkar Muralidharan , Georges Natsoulis , John Bell , Hanlee Ji , Nancy R. Zhang

Not all data in a typical training set help with generalization; some samples can be overly ambiguous or outrightly mislabeled. This paper introduces a new method to identify such samples and mitigate their impact when training neural…

Machine Learning · Computer Science 2020-12-24 Geoff Pleiss , Tianyi Zhang , Ethan R. Elenberg , Kilian Q. Weinberger

Several modern genomic technologies, such as DNA-Methylation arrays, measure spatially registered probes that number in the hundreds of thousands across multiplechromosomes. The measured probes are by themselves less interesting…

Applications · Statistics 2016-11-16 John Nagorski , Genevera I. Allen

The widespread adoption of large language models (LLMs) has made it difficult to distinguish human writing from machine-produced text in many real applications. Detectors that were effective for one generation of models tend to degrade when…

Computation and Language · Computer Science 2025-12-09 Sepyan Purnama Kristanto , Lutfi Hakim , Dianni Yusuf

The analysis of polycrystalline materials benefits greatly from accurate quantitative descriptions of their grain structures. Laguerre tessellations approximate such grain structures very well. However, it is a quite challenging problem to…

Materials Science · Physics 2016-02-03 Aaron Spettl , Tim Brereton , Qibin Duan , Thomas Werz , Carl E. Krill , Dirk P. Kroese , Volker Schmidt

High throughput genome sequencing technologies such as RNA-Seq and Microarray have the potential to transform clinical decision making and biomedical research by enabling high-throughput measurements of the genome at a granular level.…

Metagenomic binning is an essential task in analyzing metagenomic sequence datasets. To analyze structure or function of microbial communities from environmental samples, metagenomic sequence fragments are assigned to their taxonomic…

Quantitative Methods · Quantitative Biology 2016-04-12 Yunan Luo , Jianyang Zeng , Bonnie Berger , Jian Peng

Reconciling gene trees with a species tree is a fundamental problem to understand the evolution of gene families. Many existing approaches reconcile each gene tree independently. However, it is well-known that the evolution of gene families…

Populations and Evolution · Quantitative Biology 2018-06-12 Riccardo Dondi , Manuel Lafond , Celine Scornavacca

While mislabeled or ambiguously-labeled samples in the training set could negatively affect the performance of deep models, diagnosing the dataset and identifying mislabeled samples helps to improve the generalization power. Training…

Computer Vision and Pattern Recognition · Computer Science 2022-12-21 Qingrui Jia , Xuhong Li , Lei Yu , Jiang Bian , Penghao Zhao , Shupeng Li , Haoyi Xiong , Dejing Dou

Segmental duplications (SDs), or low-copy repeats (LCR), are segments of DNA greater than 1 Kbp with high sequence identity that are copied to other regions of the genome. SDs are among the most important sources of evolution, a common…

Data Structures and Algorithms · Computer Science 2018-09-25 Ibrahim Numanagić , Alim S. Gökkaya , Lillian Zhang , Bonnie Berger , Can Alkan , Faraz Hach

A quest to determine the complete sequence of a human DNA from telomere to telomere started three decades ago and was finally completed in 2021. This accomplishment was a result of a tremendous effort of numerous experts who engineered…

Genomics · Quantitative Biology 2022-06-03 Lovro Vrček , Xavier Bresson , Thomas Laurent , Martin Schmitz , Mile Šikić

Labeling of DNA molecules is a fundamental technique for DNA visualization and analysis. This process was mathematically modeled in [1], where the received sequence indicates the positions of the used labels. In this work, we develop error…

Information Theory · Computer Science 2025-11-04 Dganit Hanania , Eitan Yaakobi

Huge numbers of short reads are being generated for mapping back to the genome to discover the frequency of transcripts, miRNAs, DNAase hypersensitive sites, FAIRE regions, nucleosome occupancy, etc. Since these reads are typically short…

Genomics · Quantitative Biology 2009-12-31 Peter J. Waddell , Timothy Herston

We present Meraculous2, an update to the Meraculous short-read assembler that includes (1) handling of allelic variation using "bubble" structures within the de Bruijn graph, (2) improved gap closing, and (3) an improved scaffolding…

Data Structures and Algorithms · Computer Science 2017-11-09 Jarrod A. Chapman , Isaac Y. Ho , Eugene Goltsman , Daniel S. Rokhsar