English
Related papers

Related papers: Towards Better Understanding of Artifacts in Varia…

200 papers

The growing use of convolutional neural networks (CNN) for a broad range of visual tasks, including tasks involving fine details, raises the problem of applying such networks to a large field of view, since the amount of computations…

Computer Vision and Pattern Recognition · Computer Science 2018-04-11 Hadar Gorodissky , Daniel Harari , Shimon Ullman

Background: Regions with copy number variations (in germline cells) or copy number alteration (in somatic cells) are of great interest for human disease gene mapping and cancer studies. They represent a new type of mutation and are…

Genomics · Quantitative Biology 2012-05-07 Wentian Li , Annette Lee , Peter K Gregersen

Replication helps ensure that a genotype-phenotype association observed in a genome-wide association (GWA) study represents a credible association and is not a chance finding or an artifact due to uncontrolled biases. We discuss…

Methodology · Statistics 2010-10-26 Peter Kraft , Eleftheria Zeggini , John P. A. Ioannidis

Late diagnosis and high costs are key factors that negatively impact the care of cancer patients worldwide. Although the availability of biological markers for the diagnosis of cancer type is increasing, costs and reliability of tests…

Machine Learning · Computer Science 2019-08-20 Sterling Ramroach , Melford John , Ajay Joshi

We construct genomic predictors for heritable and extremely complex human quantitative traits (height, heel bone density, and educational attainment) using modern methods in high dimensional statistics (i.e., machine learning). Replication…

Genomics · Quantitative Biology 2017-09-20 Louis Lello , Steven G. Avery , Laurent Tellier , Ana Vazquez , Gustavo de los Campos , Stephen D. H. Hsu

In this paper, we solve the sample shortage problem in the human parsing task. We begin with the self-learning strategy, which generates pseudo-labels for unlabeled data to retrain the model. However, directly using noisy pseudo-labels will…

Computer Vision and Pattern Recognition · Computer Science 2020-04-20 Tao Li , Zhiyuan Liang , Sanyuan Zhao , Jiahao Gong , Jianbing Shen

This paper studies the haplotype assembly problem from an information theoretic perspective. A haplotype is a sequence of nucleotide bases on a chromosome, often conveniently represented by a binary string, that differ from the bases in the…

Information Theory · Computer Science 2014-05-13 Hongbo Si , Haris Vikalo , Sriram Vishwanath

Structural variants compose the majority of human genetic variation, but are difficult to assess using current genomic sequencing technologies. Optical mapping technologies, which measure the size of chromosomal fragments between labeled…

Quantitative Methods · Quantitative Biology 2019-10-10 Weiwei Li , Jan Hannig , Corbin Jones

Standard resampling ratios (e.g., $\alpha \approx 0.632$) are widely used as default baselines in ensemble learning for three decades. However, how these ratios interact with a base learner's intrinsic functional complexity in finite…

Machine Learning · Computer Science 2026-04-14 Ye Su , Mingrui Ye , Yining Wang , Jipeng Guo , Yong Liu

While most current high-throughput DNA sequencing technologies generate short reads with low error rates, emerging sequencing technologies generate long reads with high error rates. A basic question of interest is the tradeoff between read…

Information Theory · Computer Science 2015-01-27 Ilan Shomorony , Thomas Courtade , David Tse

The development of single-cell technologies provides the opportunity to identify new cellular states and reconstruct novel cell-to-cell relationships. Applications range from understanding the transcriptional and epigenetic processes…

Quantitative Methods · Quantitative Biology 2018-10-11 Luis Aparicio , Mykola Bordyuh , Andrew J. Blumberg , Raul Rabadan

Effective and reliable data retrieval is critical for the feasibility of DNA storage, and the development of random access efficiency plays a key role in its practicality and reliability. In this paper, we study the Random Access Problem,…

Information Theory · Computer Science 2025-10-10 Anina Gruica , Maria Montanucci , Ferdinando Zullo

We consider applying Bayesian Variable Selection Regression, or BVSR, to genome-wide association studies and similar large-scale regression problems. Currently, typical genome-wide association studies measure hundreds of thousands, or…

Applications · Statistics 2011-10-28 Yongtao Guan , Matthew Stephens

The ability to store data in the DNA of a living organism has applications in a variety of areas including synthetic biology and watermarking of patented genetically-modified organisms. Data stored in this medium is subject to errors…

Information Theory · Computer Science 2016-11-17 Siddharth Jain , Farzad Farnoud , Moshe Schwartz , Jehoshua Bruck

Histo-genomic multi-modal methods have recently emerged as a powerful paradigm, demonstrating significant potential for improving cancer prognosis. However, genome sequencing, unlike histopathology imaging, is still not widely accessible in…

Image and Video Processing · Electrical Eng. & Systems 2024-03-19 Zhikang Wang , Yumeng Zhang , Yingxue Xu , Seiya Imoto , Hao Chen , Jiangning Song

Recent advances in genomic sequencing technology have resulted in an abundance of genome sequence data. Despite the progress in interpreting those data, there remains a broad scope for their translation into clinical and societal benefits.…

Genomics · Quantitative Biology 2021-12-13 Abhinav Jain , Greg Slabaugh , Deepti Gurdasani

In this paper we study error-correcting codes for the storage of data in synthetic deoxyribonucleic acid (DNA). We investigate a storage model where a data set is represented by an unordered set of $M$ sequences, each of length $L$. Errors…

Information Theory · Computer Science 2020-02-13 Andreas Lenz , Paul H. Siegel , Antonia Wachter-Zeh , Eitan Yaakobi

Genotype imputation enables dense variant coverage for genome-wide association and risk-prediction studies, yet conventional reference-panel methods remain limited by ancestry bias and reduced rare-variant accuracy. We present Genotype…

Given a set of aligned sequences of independent noisy observations, we are concerned with detecting intervals where the mean values of the observations change simultaneously in a subset of the sequences. The intervals of changed means are…

Applications · Statistics 2011-08-17 David Siegmund , Benjamin Yakir , Nancy R. Zhang

Understanding genetic variation, e.g., through mutations, in organisms is crucial to unravel their effects on the environment and human health. A fundamental characterization can be obtained by solving the haplotype assembly problem, which…

Genomics · Quantitative Biology 2022-10-25 Hansheng Xue , Vaibhav Rajan , Yu Lin