English
Related papers

Related papers: Improving sequence-based genotype calls with linka…

200 papers

Although prospective logistic regression is the standard method of analysis for case-control data, it has been recently noted that in genetic epidemiologic studies one can use the ``retrospective'' likelihood to gain major power by…

Methodology · Statistics 2010-10-25 Nilanjan Chatterjee , Yi-Hau Chen , Sheng Luo , Raymond J. Carroll

Motivation: Predicting gene-disease associations (GDAs) is the problem to determine which gene is associated with a disease. GDA prediction can be framed as a ranking problem where genes are ranked for a query disease, based on features…

Quantitative Methods · Quantitative Biology 2026-02-03 Fernando Zhapa-Camacho , Robert Hoehndorf

Certifiable, adaptive uncertainty estimates for unknown quantities are an essential ingredient of sequential decision-making algorithms. Standard approaches rely on problem-dependent concentration results and are limited to a specific…

Machine Learning · Computer Science 2023-11-09 Nicolas Emmenegger , Mojmír Mutný , Andreas Krause

DNA samples are often pooled, either by experimental design, or because the sample itself is a mixture. For example, when population allele frequencies are of primary interest, individual samples may be pooled together to lower the cost of…

Quantitative Methods · Quantitative Biology 2013-02-07 Darren Kessner , Tom Turner , John Novembre

When searching for gene pathways leading to specific disease outcomes, additional information on gene characteristics is often available that may facilitate to differentiate genes related to the disease from irrelevant background when…

Machine Learning · Statistics 2018-09-10 Yunpeng Zhao , Qing Pan , Chengan Du

DNA databases are widely used in forensic science to identify unknown offenders. When no exact match is found, familial DNA searches can help by identifying first-degree relatives using likelihood ratios. If multiple subpopulations are…

Applications · Statistics 2025-12-08 Monchai Kooakachai , Tiwakorn Chapalee , Chairat Thitiyan , Patsaya Jumnongwut

With the advance of high-throughput sequencing technologies, it has become feasible to investigate the influence of the entire spectrum of sequencing variations on complex human diseases. Although association studies utilizing the new…

Methodology · Statistics 2025-08-19 Ming Li , Zihuai He , Min Zhang , Xiaowei Zhan , Changshuai Wei , Robert C Elston , Qing Lu

A computational challenge to validate the candidate disease genes identified in a high-throughput genomic study is to elucidate the associations between the set of candidate genes and disease phenotypes. The conventional gene set enrichment…

Genomics · Quantitative Biology 2011-02-22 TaeHyun Hwang , Wei Zhang , Maoqiang Xie , Rui Kuang

Complete genome sequences contain valuable information about natural selection, but extracting this information for short, widely scattered noncoding elements remains a challenging problem. Here we introduce a new computational method for…

Genomics · Quantitative Biology 2015-03-19 Ilan Gronau , Leonardo Arbiza , Jaaved Mohammed , Adam Siepel

Motivated by genome-wide association studies, we consider a standard linear model with one additional random effect in situations where many predictors have been collected on the same subjects and each predictor is analyzed separately.…

Applications · Statistics 2013-04-24 Matti Pirinen , Peter Donnelly , Chris C. A. Spencer

Deep neural network-based architectures give promising results in various domains including pattern recognition. Finding the optimal combination of the hyper-parameters of such a large-sized architecture is tedious and requires a large…

Computer Vision and Pattern Recognition · Computer Science 2020-03-17 Animesh Singh , Sandip Saha , Ritesh Sarkhel , Mahantapas Kundu , Mita Nasipuri , Nibaran Das

The alignment of biological sequences such as DNA, RNA, and proteins, is one of the basic tools that allow to detect evolutionary patterns, as well as functional/structural characterizations between homologous sequences in different…

Quantitative Methods · Quantitative Biology 2023-05-01 Louise Budzynski , Andrea Pagnani

Improving existing widely-adopted prediction models is often a more efficient and robust way towards progress than training new models from scratch. Existing models may (a) incorporate complex mechanistic knowledge, (b) leverage proprietary…

Gene set collections are a common ground to study the enrichment of genes for specific phenotypic traits. Gene set enrichment analysis aims to identify genes that are over-represented in gene sets collections and might be associated with a…

Genomics · Quantitative Biology 2022-07-26 Chiara Balestra , Carlo Maj , Emmanuel Mueller , Andreas Mayr

Metagenomics characterizes the taxonomic diversity of microbial communities by sequencing DNA directly from an environmental sample. One of the main challenges in metagenomics data analysis is the binning step, where each sequenced read is…

Quantitative Methods · Quantitative Biology 2015-05-27 Kévin Vervier , Pierre Mahé , Maud Tournoud , Jean-Baptiste Veyrieras , Jean-Philippe Vert

Demographic models built from genetic data play important roles in illuminating prehistorical events and serving as null models in genome scans for selection. We introduce an inference method based on the joint frequency spectrum of genetic…

Populations and Evolution · Quantitative Biology 2010-05-10 Ryan N. Gutenkunst , Ryan D. Hernandez , Scott H. Williamson , Carlos D. Bustamante

We introduce Phen-Gen, a method which combines patient disease symptoms and sequencing data with prior domain knowledge to identify the causative gene(s) for rare disorders.

Genomics · Quantitative Biology 2015-03-02 Asif Javed , Saloni Agrawal , Pauline C. Ng

To date, efforts to produce high-quality polygenic risk scores from genome-wide studies of common disease have focused on estimating and aggregating the effects of multiple SNPs. Here we propose a novel statistical approach for genetic risk…

Quantitative Methods · Quantitative Biology 2014-05-13 David Golan , Saharon Rosset

Motivation: Next generation methods of DNA sequencing produce relatively high rate of reading errors, which interfere with de novo genome assembly of newly sequenced organisms and particularly affect the quality of SNP detection important…

Genomics · Quantitative Biology 2019-07-31 Oleg Fokin , Anastasia Bakulina , Igor Seledtsov , Victor Solovyev

We investigate saddlepoint approximations applied to the score test statistic in genome-wide association studies with binary phenotypes. The inaccuracy in the normal approximation of the score test statistic increases with increasing sample…