English
Related papers

Related papers: Finding easy regions for short-read variant callin…

200 papers

Background: Structural variants (SVs) are genomic differences $\ge$50 bp in length. They remain challenging to detect even with long sequence reads, and the sources of these difficulties are not well quantified. Results: We identified 35.4…

Genomics · Quantitative Biology 2025-09-30 Qian Qin , Heng Li

Motivation: Whole-genome high-coverage sequencing has been widely used for personal and cancer genomics as well as in various research areas. However, in the lack of an unbiased whole-genome truth set, the global error rate of variant calls…

Genomics · Quantitative Biology 2018-07-27 Heng Li

Background: Several sources of noise obfuscate the identification of single nucleotide variation (SNV) in next generation sequencing data. For instance, errors may be introduced during library construction and sequencing steps. In addition,…

Genomics · Quantitative Biology 2015-03-05 Steve Hoffmann , Peter F. Stadler , Korbinian Strimmer

Variant calling, the problem of estimating whether a position in a DNA sequence differs from a reference sequence, given noisy, redundant, overlapping short sequences that cover that position, is fundamental to genomics. We propose a deep…

Genomics · Quantitative Biology 2020-03-17 Nikolai Yakovenko , Avantika Lal , Johnny Israeli , Bryan Catanzaro

DNA read mapping is a ubiquitous task in bioinformatics, and many tools have been developed to solve the read mapping problem. However, there are two trends that are changing the landscape of readmapping: First, new sequencing technologies…

Genomics · Quantitative Biology 2017-02-09 Jens Quedenfeld , Sven Rahmann

De novo genome assembly is challenging in highly repetitive regions; however, reference-guided assemblers often suffer from bias. We propose a framework for pangenome-guided sequence assembly, which can resolve short-read data in complex…

Quantum Physics · Physics 2026-02-11 Josh Cudby , James Bonfield , Chenxi Zhou , Richard Durbin , Sergii Strelchuk

Confidence region prediction is a practically useful extension to the commonly studied pattern recognition problem. Instead of predicting a single label, the constraint is relaxed to allow prediction of a subset of labels given a desired…

Machine Learning · Computer Science 2024-05-27 David Lindsay , Sian Lindsay

Whole and targeted sequencing of human genomes is a promising, increasingly feasible tool for discovering genetic contributions to risk of complex diseases. A key step is calling an individual's genotype from the multiple aligned short read…

Applications · Statistics 2012-06-29 Baiyu Zhou , Alice S. Whittemore

Cheap high-throughput DNA sequencing may soon become routine not only for human genomes but also for practically anything requiring the identification of living organisms from their DNA: tracking of infectious agents, control of food…

Genomics · Quantitative Biology 2014-03-05 Laurent Gautier , Ole Lund

Research on the localization of the genetic basis associated with diseases or traits has been widely conducted in the last a few decades. Scan methods have been developed for region-based analysis in whole-genome association studies,…

Methodology · Statistics 2024-10-31 Wei Zhang , Fan Wang , Fang Yao

High-throughput shotgun sequence data makes it possible in principle to accurately estimate population genetic parameters without confounding by SNP ascertainment bias. One such statistic of interest is the proportion of heterozygous sites…

Populations and Evolution · Quantitative Biology 2012-12-18 Katarzyna Bryc , Nick Patterson , David Reich

Pangenome variation graphs (PVGs) allow for the representation of genetic diversity in a more nuanced way than traditional reference-based approaches. Here we focus on how PVGs are a powerful tool for studying genetic variation in viruses,…

Genomics · Quantitative Biology 2025-06-18 Tim Downing

We show that confidence intervals in a variance component model, with asymptotically correct uniform coverage probability, can be obtained by inverting certain test-statistics based on the score for the restricted likelihood. The results…

Methodology · Statistics 2024-11-05 Yiqiao Zhang , Karl Oskar Ekvall , Aaron J. Molstad

It is increasingly common clinically for cancer specimens to be examined using techniques that identify somatic mutations. In principle these mutational profiles can be used to diagnose the tissue of origin, a critical task for the 3-5% of…

Methodology · Statistics 2020-07-14 Saptarshi Chakraborty , Colin B. Begg , Ronglai Shen

Conducting genome-wide association studies (GWAS) in copy number variation (CNV) level is a field where few people involves and little statistical progresses have been achieved, traditional methods suffer from many problems such as batch…

Methodology · Statistics 2020-11-17 Han Wang , Changhu Wang , Linjie Wu , Ruibin Xi

High read depth can be used to assemble short sequence repeats. The existing genome assemblers fail in repetitive regions of longer than average read. I propose a new algorithm for a DNA assembly which uses the relative frequency of reads…

Genomics · Quantitative Biology 2015-01-08 Robert M. Nowak

DNA sequencing to identify genetic variants is becoming increasingly valuable in clinical settings. Assessment of variants in such sequencing data is commonly implemented through Bayesian heuristic algorithms. Machine learning has shown…

Successful sequencing experiments require judicious sample selection. However, this selection must often be performed on the basis of limited preliminary data. Predicting the statistical properties of the final sample based on preliminary…

Genomics · Quantitative Biology 2014-03-17 Simon Gravel , NHLBI GO Exome Sequencing Project

Huge numbers of short reads are being generated for mapping back to the genome to discover the frequency of transcripts, miRNAs, DNAase hypersensitive sites, FAIRE regions, nucleosome occupancy, etc. Since these reads are typically short…

Genomics · Quantitative Biology 2009-12-31 Peter J. Waddell , Timothy Herston

Construction of tight confidence regions and intervals is central to statistical inference and decision making. This paper develops new theory showing minimum average volume confidence regions for categorical data. More precisely, consider…

Machine Learning · Statistics 2021-02-01 Matthew L. Malloy , Ardhendu Tripathy , Robert D. Nowak
‹ Prev 1 2 3 10 Next ›