English
Related papers

Related papers: Minimum error correction-based haplotype assembly:…

200 papers

A critical problem in the emerging high-throughput genotyping protocols is to minimize the number of polymerase chain reaction (PCR) primers required to amplify the single nucleotide polymorphism loci of interest. In this paper we study PCR…

Data Structures and Algorithms · Computer Science 2007-05-23 K. Konwar , I. Mandoiu , A. Russell , A. Shvartsman

The Constrained Minimal Supersymmetric Standard Model (CMSSM) is one of the simplest and most widely-studied supersymmetric extensions to the standard model of particle physics. Nevertheless, current data do not sufficiently constrain the…

High Energy Physics - Phenomenology · Physics 2015-03-13 Yashar Akrami , Pat Scott , Joakim Edsjö , Jan Conrad , Lars Bergström

Short tandem repeats (STRs) and single nucleotide polymorphisms (SNPs) are two kinds of commonly used markers in Y chromosome studies of forensic and population genetics. There has been increasing interest in the cost saving strategy by…

Populations and Evolution · Quantitative Biology 2013-10-22 Chuan-Chao Wang , Ling-Xiang Wang , Rukesh Shrestha , Shaoqing Wen , Manfei Zhang , Xinzhu Tong , Li Jin , Hui Li

Synthetic DNA can in principle be used for the archival storage of arbitrary data. Because errors are introduced during DNA synthesis, storage, and sequencing, an error-correcting code (ECC) is necessary for error-free recovery of the data.…

Quantitative Methods · Quantitative Biology 2018-12-05 William H. Press , John A. Hawkins

Segmental duplications (SDs), or low-copy repeats (LCR), are segments of DNA greater than 1 Kbp with high sequence identity that are copied to other regions of the genome. SDs are among the most important sources of evolution, a common…

Data Structures and Algorithms · Computer Science 2018-09-25 Ibrahim Numanagić , Alim S. Gökkaya , Lillian Zhang , Bonnie Berger , Can Alkan , Faraz Hach

Motivation: Genome-Wide Association Studies (GWAS) seek to identify causal genomic variants associated with rare human diseases. The classical statistical approach for detecting these variants is based on univariate hypothesis testing, with…

Methodology · Statistics 2018-10-22 Florent Guinot , Marie Szafranski , Christophe Ambroise , Franck Samson

Motivation: Next generation methods of DNA sequencing produce relatively high rate of reading errors, which interfere with de novo genome assembly of newly sequenced organisms and particularly affect the quality of SNP detection important…

Genomics · Quantitative Biology 2019-07-31 Oleg Fokin , Anastasia Bakulina , Igor Seledtsov , Victor Solovyev

In this work we present a flexible, probabilistic and reference-free method of error correction for high throughput DNA sequencing data. The key is to exploit the high coverage of sequencing data and model short sequence outputs as…

Information Theory · Computer Science 2013-02-04 Xin Yin , Zhao Song , Karin Dorman , Aditya Ramamoorthy

New long read sequencing technologies, like PacBio SMRT and Oxford NanoPore, can produce sequencing reads up to 50,000 bp long but with an error rate of at least 15%. Reducing the error rate is necessary for subsequent utilisation of the…

Genomics · Quantitative Biology 2021-11-18 Leena Salmela , Riku Walve , Eric Rivals , Esko Ukkonen

Sparse coding (Sc) has been studied very well as a powerful data representation method. It attempts to represent the feature vector of a data sample by reconstructing it as the sparse linear combination of some basic elements, and a $L_2$…

Machine Learning · Computer Science 2016-03-15 Mohua Zhang , Jianhua Peng , Xuejie Liu , Jim Jing-Yan Wang

The widely used genetic pleiotropic analysis of multiple phenotypes are often designed for examining the relationship between common variants and a few phenotypes. They are not suited for both high dimensional phenotypes and high…

Machine Learning · Statistics 2015-12-04 Panpan Wang , Mohammad Rahman , Li Jin , Momiao Xiong

Despite much progress over the past decade, current Single Nucleotide Polymorphism (SNP) genotyping technologies still offer an insufficient degree of multiplexing when required to handle user-selected sets of SNPs. In this paper we propose…

Data Structures and Algorithms · Computer Science 2007-05-23 Ion I. Mandoiu , Claudia Prajescu

We propose masked particle modeling (MPM) as a self-supervised method for learning generic, transferable, and reusable representations on unordered sets of inputs for use in high energy physics (HEP) scientific data. This work provides a…

High Energy Physics - Phenomenology · Physics 2024-07-12 Tobias Golling , Lukas Heinrich , Michael Kagan , Samuel Klein , Matthew Leigh , Margarita Osadchy , John Andrew Raine

Computing haplotypes from sequencing data, i.e. haplotype assembly, is an important component of molecular and population genetics problems, including interpreting the effects of genetic variation on complex traits and reconstructing…

Genomics · Quantitative Biology 2026-03-12 Marjan Hosseini , Ella Veiner , Thomas Bergendahl , Tala Yasenpoor , Zane Smith , Margaret Staton , Derek Aguiar

Accurate identification of haplotypes in sequenced human genomes can provide invaluable information about population demography and fine-scale correlations along the genome, thus empowering both population genomic and medical association…

Genomics · Quantitative Biology 2012-11-12 Fouad Zakharia , Carlos Bustamante

Analysis of genomic segments shared identical-by-descent (IBD) between individuals is fundamental to many genetic applications, from demographic inference to estimating the heritability of diseases, but IBD detection accuracy in…

Populations and Evolution · Quantitative Biology 2014-02-11 Eric Y. Durand , Nicholas Eriksson , Cory Y. McLean

Sparse autoencoders (SAEs) are widely used for interpreting language model activations. A key evaluation metric is the increase in cross-entropy loss between the original model logits and the reconstructed model logits when replacing model…

Machine Learning · Computer Science 2025-04-01 Adam Karvonen

The prevalent technique for DNA sequencing consists of two main steps: shotgun sequencing, where many randomly located fragments, called reads, are extracted from the overall sequence, followed by an assembly algorithm that aims to…

Genomics · Quantitative Biology 2016-01-28 Shirshendu Ganguly , Elchanan Mossel , Miklos Z. Racz

Statistically resolving the underlying haplotype pair for a genotype measurement is an important intermediate step in gene mapping studies, and has received much attention recently. Consequently, a variety of methods for this problem have…

Machine Learning · Computer Science 2007-10-29 Matti Kääriäinen , Niels Landwehr , Sampsa Lappalainen , Taneli Mielikäinen

Error correction of sequenced reads remains a difficult task, especially in single-cell sequencing projects with extremely non-uniform coverage. While existing error correction tools designed for standard (multi-cell) sequencing data…

Quantitative Methods · Quantitative Biology 2013-01-31 Sergey I. Nikolenko , Anton I. Korobeynikov , Max A. Alekseyev