English
Related papers

Related papers: Estimating heterozygosity from a low-coverage geno…

200 papers

DNA storage is now being considered as a new archival storage method for its durability and high information density, but still facing some challenges like high costs and low throughput. By reducing sequencing sample size for decoding…

Information Theory · Computer Science 2025-04-22 Ruiying Cao , Xin Chen

This issue includes six articles that develop and apply statistical methods for the analysis of gene sequencing data of different types. The methods are tailored to the different data types and, in each case, lead to biological insights not…

Applications · Statistics 2012-06-29 Karen Kafadar

Human populations have experienced dramatic growth since the Neolithic revolution. Recent studies that sequenced a very large number of individuals observed an extreme excess of rare variants, and provided clear evidence of recent rapid…

Populations and Evolution · Quantitative Biology 2015-06-17 Elodie Gazave , Li Ma , Diana Chang , Alex Coventry , Feng Gao , Donna Muzny , Eric Boerwinkle , Richard Gibbs , Charles F. Sing , Andrew G. Clark , Alon Keinan

Diagnosis and risk stratification of cancer and many other diseases require the detection of genomic breakpoints as a prerequisite of calling copy number alterations (CNA). This, however, is still challenging and requires time-consuming…

DNA sequencing to identify genetic variants is becoming increasingly valuable in clinical settings. Assessment of variants in such sequencing data is commonly implemented through Bayesian heuristic algorithms. Machine learning has shown…

The computational problem of inferring the full haplotype of a cell starting from read sequencing data is known as haplotype assembly, and consists in assigning all heterozygous Single Nucleotide Polymorphisms (SNPs) to exactly one of the…

Large-scale statistical analysis of data sets associated with genome sequences plays an important role in modern biology. A key component of such statistical analyses is the computation of $p$-values and confidence bounds for statistics…

Applications · Statistics 2011-01-06 Peter J. Bickel , Nathan Boley , James B. Brown , Haiyan Huang , Nancy R. Zhang

It is increasingly common clinically for cancer specimens to be examined using techniques that identify somatic mutations. In principle these mutational profiles can be used to diagnose the tissue of origin, a critical task for the 3-5% of…

Methodology · Statistics 2020-07-14 Saptarshi Chakraborty , Colin B. Begg , Ronglai Shen

Statistically resolving the underlying haplotype pair for a genotype measurement is an important intermediate step in gene mapping studies, and has received much attention recently. Consequently, a variety of methods for this problem have…

Machine Learning · Computer Science 2007-10-29 Matti Kääriäinen , Niels Landwehr , Sampsa Lappalainen , Taneli Mielikäinen

Heterogeneity is a fundamental characteristic of cancer. To accommodate heterogeneity, subgroup identification has been extensively studied and broadly categorized into unsupervised and supervised analysis. Compared to unsupervised…

Methodology · Statistics 2026-02-25 Xing Qin , Xu Liu , Shuangge Ma , Mengyun Wu

We develop statistically based methods to detect single nucleotide DNA mutations in next generation sequencing data. Sequencing generates counts of the number of times each base was observed at hundreds of thousands to billions of genome…

Applications · Statistics 2012-10-01 Omkar Muralidharan , Georges Natsoulis , John Bell , Hanlee Ji , Nancy R. Zhang

Recently, ultra high-throughput sequencing of RNA (RNA-Seq) has been developed as an approach for analysis of gene expression. By obtaining tens or even hundreds of millions of reads of transcribed sequences, an RNA-Seq experiment can offer…

Methodology · Statistics 2011-06-17 Julia Salzman , Hui Jiang , Wing Hung Wong

The distribution of bases spacing in human genome was investigated. An analysis of the frequency of occurrence in the human genome of different sequence lengths flanked by one type of nucleotide was carried out showing that the distribution…

Other Quantitative Biology · Quantitative Biology 2019-04-18 Andrzej Z. Górski , Monika Piwowar

Motivated by applications in neuroanatomy, we propose a novel methodology for estimating the heritability which corresponds to the proportion of phenotypic variance which can be explained by genetic factors. Estimating this quantity for…

Statistics Theory · Mathematics 2016-06-09 Anna Bonnet , Céline Lévy-Leduc , Elisabeth Gassiat , Roberto Toro , Thomas Bourgeron

We present a framework for the design of optimal assembly algorithms for shotgun sequencing under the criterion of complete reconstruction. We derive a lower bound on the read length and the coverage depth required for reconstruction in…

Genomics · Quantitative Biology 2013-02-20 Guy Bresler , Ma'ayan Bresler , David Tse

We consider the problem of detecting and estimating the strength of association between a trait of interest and alleles or haplotypes in a small genomic region (e.g. a gene or a gene complex), when no direct information on that region is…

Applications · Statistics 2008-04-11 Rodrigo Labouriau , Poul Sørensen , Helle R. Juul-Madsen

In the effort to aid cytologic diagnostics by establishing automatic single cell screening using high throughput digital holographic microscopy for clinical studies thousands of images and millions of cells are captured. The bottleneck lies…

Image and Video Processing · Electrical Eng. & Systems 2024-10-28 Julia Sistermanns , Ellen Emken , Gregor Weirich , Oliver Hayden , Wolfgang Utschick

Covariance matrix estimation is a fundamental statistical task in many applications, but the sample covariance matrix is sub-optimal when the sample size is comparable to or less than the number of features. Such high-dimensional settings…

Methodology · Statistics 2022-06-06 Huiqin Xin , Sihai Dave Zhao

Through genome-wide association studies (GWAS), disease susceptible genetic variables can be identified by comparing the genetic data of individuals with and without a specific disease. However, the discovery of these associations poses a…

Machine Learning · Computer Science 2023-08-15 Zhendong Sha , Yuanzhu Chen , Ting Hu

Locating recombination hotspots in genomic data is an important but difficult task. Current methods frequently rely on estimating complicated models at high computational cost. In this paper we develop an extremely fast, scalable method for…

Applications · Statistics 2015-12-08 Jordan Rodu , Shane T. Jensen