English
Related papers

Related papers: Ancestral inference from haplotypes and mutations

200 papers

In this paper, we consider the problem of numerical investigation of the counting statistics for a class of one-dimensional systems. Importance sampling, the cornerstone technique usually implemented for such problems, critically hinges on…

Statistical Mechanics · Physics 2024-08-12 Ivan N. Burenev , Satya N. Majumdar , Alberto Rosso

This article presents new methodology for sample-based Bayesian inference when data are partitioned and communication between the parts is expensive, as arises by necessity in the context of "big data" or by choice in order to take…

Methodology · Statistics 2022-11-01 Marc Box

Much of the on-going statistical analysis of DNA sequences is focused on the estimation of characteristics of coding and non-coding regions that would possibly allow discrimination of these regions. In the current approach, we concentrate…

Genomics · Quantitative Biology 2009-11-10 D. Kugiumtzis , A. Provata

In many epidemiological contexts, disease occurrences and their rates are naturally modelled by counting processes and their intensities, allowing an analysis based on martingale methods. These methods lend themselves to extensions of…

Statistics Theory · Mathematics 2007-06-13 Larry Goldstein , Bryan Langholz

The sample frequency spectrum of a segregating site is the probability distribution of a sample of alleles from a genetic locus, conditional on observing the sample to have more than one clearly different phenotypes. We present a model for…

Probability · Mathematics 2014-05-13 Arka Bhattacharya

Recent improvements in high-throughput genotyping and sequencing technologies have afforded the collection of massive, genome-wide datasets of DNA information from hundreds of thousands of individuals. These datasets, in turn, provide…

Populations and Evolution · Quantitative Biology 2014-12-19 Pier Francesco Palamara

We develop a Bayesian approach for selecting the model which is the most supported by the data within a class of marginal models for categorical variables formulated through equality and/or inequality constraints on generalised logits…

Statistics Theory · Mathematics 2012-02-21 Francesco Bartolucci , Luisa Scaccia , Alessio Farcomeni

The Bayesian evidence, crucial ingredient for model selection, is arguably the most important quantity in Bayesian data analysis: at the same time, however, it is also one of the most difficult to compute. In this paper we present a…

Methodology · Statistics 2024-05-14 Stefano Rinaldi , Gabriele Demasi , Walter Del Pozzo , Otto A. Hannuksela

This paper presents and analyzes an approach to cluster-based inference for dependent data. The primary setting considered here is with spatially indexed data in which the dependence structure of observed random variables is characterized…

Statistics Theory · Mathematics 2022-11-16 Jianfei Cao , Christian Hansen , Damian Kozbur , Lucciano Villacorta

Sampling from a multimodal distribution is a fundamental and challenging problem in computational science and statistics. Among various approaches proposed for this task, one popular method is Annealed Importance Sampling (AIS). In this…

Computation · Statistics 2024-11-07 Haoxuan Chen , Lexing Ying

In considering evolution of transcribed regions, regulatory modules, and other genomic loci of interest, we are often faced with a situation in which the number of allelic states greatly exceeds the population size. In this limit, the…

Populations and Evolution · Quantitative Biology 2016-07-27 Pavel Khromov , Constantin D. Malliaris , Alexandre V. Morozov

Genetic data obtained on population samples convey information about their evolutionary history. Inference methods can extract this information (at least partially) but they require sophisticated statistical techniques that have been made…

When can reliable inference be drawn in the "Big Data" context? This paper presents a framework for answering this fundamental question in the context of correlation mining, with implications for general large scale inference. In large…

Statistics Theory · Mathematics 2015-05-19 Alfred O. Hero , Bala Rajaratnam

The ancestral sequence reconstruction problem is the inference, back in time, of the properties of common sequence ancestors from measured properties of contemporary populations. Standard algorithms for this problem assume independent…

Disordered Systems and Neural Networks · Physics 2022-02-09 Edwin Rodríguez Horta , Alejandro Lage-Castellanos , Roberto Mulet

We consider an evolving system for which a sequence of observations is being made, with each observation revealing additional information about current and past states of the system. We suppose each observation is made without error, but…

Computation · Statistics 2021-03-10 Valentina Di Marco , Jonathan Keith

In phylogenetic inference one is interested in obtaining samples from the posterior distribution over the tree space on the basis of some observed DNA sequence data. The challenge is to obtain samples from this target distribution without…

Statistics Theory · Mathematics 2007-06-13 Raazesh Sainudiin , Thomas York

Estimating feature importance is a significant aspect of explaining data-based models. Besides explaining the model itself, an equally relevant question is which features are important in the underlying data generating process. We present a…

Machine Learning · Computer Science 2021-09-21 Pål Vegard Johnsen , Inga Strümke , Signe Riemer-Sørensen , Andrew Thomas DeWan , Mette Langaas

In population genetics, there is often interest in inferring selection coefficients. This task becomes more challenging if multiple linked selected loci are considered simultaneously. For such a situation, we propose a novel generalized…

Methodology · Statistics 2025-12-17 Ritabrata Dutta , Yuehao Xu , Sherman Khoo , Francesca Basini , Andreas Futschik

Motivated by a non-random but clustered distribution of SNPs, we introduce a phenomenological model to account for the clustering properties of SNPs in the human genome. The phenomenological model is based on a preferential mutation to the…

Genomics · Quantitative Biology 2016-05-24 Chang-Yong Lee

The sample frequency spectrum (SFS) of DNA sequences from a collection of individuals is a summary statistic which is commonly used for parametric inference in population genetics. Despite the popularity of SFS-based inference methods,…

Populations and Evolution · Quantitative Biology 2015-06-24 Jonathan Terhorst , Yun S. Song