English
Related papers

Related papers: Predicting discovery rates of genomic features

200 papers

Prediction intervals in supervised Machine Learning bound the region where the true outputs of new samples may fall. They are necessary in the task of separating reliable predictions of a trained model from near random guesses, minimizing…

Machine Learning · Computer Science 2019-12-20 Anton Akusok , Yoan Miche , Kaj-Mikael Björk , Amaury Lendasse

A simple analytical framework to study the molecular quasispecies evolution of finite populations is proposed, in which the population is assumed to be a random combination of the constiyuent molecules in each generation,i.e., linkage…

Statistical Mechanics · Physics 2016-08-31 Domingos Alves , J. F. Fontanari

Computational methods for discovering patterns of local correlations in sequences are important in computational biology. Here we show how to determine the optimal partitioning of aligned sequences into non-overlapping segments such that…

Computational Engineering, Finance, and Science · Computer Science 2012-06-26 Joseph Bockhorst , Nebojsa Jojic

Adaptation to local environments often occurs through natural selection acting on a large number of loci, each having a weak phenotypic effect. One way to detect these loci is to identify genetic polymorphisms that exhibit high correlation…

Populations and Evolution · Quantitative Biology 2015-03-20 Eric Frichot , Sean Schoville , Guillaume Bouchard , Olivier François

We investigate Bayesian predictive inference for finite population quantities when there are unequal probabilities of selection. Only limited information about the sample design is available; i.e., only the first-order selection…

Methodology · Statistics 2018-04-10 Junheng Ma , Joe Sedransk , Balgobin Nandram , Lu Chen

Motivation: Most existing methods for DNA sequence analysis rely on accurate sequences or genotypes. However, in applications of the next-generation sequencing (NGS), accurate genotypes may not be easily obtained (e.g. multi-sample…

Genomics · Quantitative Biology 2013-03-19 Heng Li

The distribution of genetic polymorphisms in a population contains information about the mutation rate and the strength of natural selection at a locus. Here, we show that the Poisson Random Field (PRF) method of population-genetic…

Populations and Evolution · Quantitative Biology 2007-07-18 Michael M Desai , Joshua B. Plotkin

We present a coherent Bayesian framework for selection of the most likely model from the five genetic models (genotypic, additive, dominant, co-dominant, and recessive) commonly used in genetic association studies. The approach uses a…

Methodology · Statistics 2015-04-22 Harold Bae , Thomas Perls , Martin Steinberg , Paola Sebastiani

Whole and targeted sequencing of human genomes is a promising, increasingly feasible tool for discovering genetic contributions to risk of complex diseases. A key step is calling an individual's genotype from the multiple aligned short read…

Applications · Statistics 2012-06-29 Baiyu Zhou , Alice S. Whittemore

The multivariate hypergeometric distribution describes sampling without replacement from a discrete population of elements divided into multiple categories. Addressing a gap in the literature, we tackle the challenge of estimating discrete…

Machine Learning · Computer Science 2024-06-11 Liam Hodgson , Danilo Bzdok

Genomic datasets generated with massively parallel sequencing methods have the potential to propel systematics in new and exciting directions, but selecting appropriate markers and methods is not straightforward. We applied two approaches…

Genomics · Quantitative Biology 2017-03-28 Michael G. Harvey , Brian Tilston Smith , Travis C. Glenn , Brant C. Faircloth , Robb T. Brumfield

Pre-trained models have been successful in many protein engineering tasks. Most notably, sequence-based models have achieved state-of-the-art performance on protein fitness prediction while structure-based models have been used…

Machine Learning · Computer Science 2023-07-25 Antonia Boca , Simon Mathis

The sample frequency spectrum (SFS) of DNA sequences from a collection of individuals is a summary statistic which is commonly used for parametric inference in population genetics. Despite the popularity of SFS-based inference methods,…

Populations and Evolution · Quantitative Biology 2015-06-24 Jonathan Terhorst , Yun S. Song

Background: With the fast development of next generation sequencing technologies, increasing numbers of genomes are being de novo sequenced and assembled. However, most are in fragmental and incomplete draft status, and thus it is often…

Genomics · Quantitative Biology 2020-02-28 Binghang Liu , Yujian Shi , Jianying Yuan , Xuesong Hu , Hao Zhang , Nan Li , Zhenyu Li , Yanxiang Chen , Desheng Mu , Wei Fan

Identifying drivers of complex traits from the noisy signals of genetic variation obtained from high throughput genome sequencing technologies is a central challenge faced by human geneticists today. We hypothesize that the variants…

Populations and Evolution · Quantitative Biology 2013-06-18 M. Cyrus Maher , Lawrence H. Uricchio , Dara G. Torgerson , Ryan D. Hernandez

With the recent advances in DNA sequencing, it is now possible to have complete genomes of individuals sequenced and assembled. This rich and focused genotype information can be used to do different population-wide studies, now first time…

Data Structures and Algorithms · Computer Science 2011-09-08 Jouni Sirén , Niko Välimäki , Veli Mäkinen

The impact of predictive algorithms on people's lives and livelihoods has been noted in medicine, criminal justice, finance, hiring and admissions. Most of these algorithms are developed using data and human capital from highly developed…

Machine Learning · Computer Science 2021-03-30 Xingyu Li , Difan Song , Miaozhe Han , Yu Zhang , Rene F. Kizilcec

Recently-developed genotype imputation methods are a powerful tool for detecting untyped genetic variants that affect disease susceptibility in genetic association studies. However, existing imputation methods require individual-level…

Applications · Statistics 2010-11-15 Xiaoquan Wen , Matthew Stephens

An explosion of high-throughput DNA sequencing in the past decade has led to a surge of interest in population-scale inference with whole-genome data. Recent work in population genetics has centered on designing inference methods for…

Machine Learning · Computer Science 2018-11-07 Jeffrey Chan , Valerio Perrone , Jeffrey P. Spence , Paul A. Jenkins , Sara Mathieson , Yun S. Song

We propose an order index, phi, which quantifies the notion of ``life at the edge of chaos'' when applied to genome sequences. It maps genomes to a number from 0 (random and of infinite length) to 1 (fully ordered) and applies regardless of…

Genomics · Quantitative Biology 2007-08-14 Sing-Guan Kong , Hong-Da Chen , Wen-Lang Fan , Jan Wigger , Andrew Torda , HC Lee