Related papers: Statistical tools for seed bank detection
The method of multivariable Mendelian randomization uses genetic variants to instrument multiple exposures, to estimate the effect that a given exposure has on an outcome conditional on all other exposures included in a linear model.…
In this paper we propose a method and discuss its computational implementation as an integrated tool for the analysis of viral genetic diversity on data generated by high-throughput sequencing. Most methods for viral diversity estimation…
In phylogenomics, species-tree methods must contend with two major sources of noise; stochastic gene-tree variation under the multispecies coalescent model (MSC) and finite-sequence substitutional noise. Fast agglomerative methods such as…
We consider species tree estimation under a standard stochastic model of gene tree evolution that incorporates incomplete lineage sorting (as modeled by a coalescent process) and gene duplication and loss (as modeled by a branching…
Phylogenetic species trees typically represent the speciation history as a bifurcating tree. Speciation events that simultaneously create more than two descendants, thereby creating polytomies in the phylogeny, are possible. Moreover, the…
Kingman's coalescent is a random tree that arises from classical population genetic models such as the Moran model. The individuals alive in these models correspond to the leaves in the tree and the following two laws of large numbers…
When an advantageous mutation occurs in a population, the favorable allele may spread to the entire population in a short time, an event known as a selective sweep. As a result, when we sample $n$ individuals from a population and trace…
The multispecies coalescent model describes the generation of gene trees from a rooted metric species tree, and thus provides a framework for the inference of species trees from sampled gene trees. We prove that the STAR method of Liu et…
Consider a random permutation of $\{1, \ldots, \lfloor n^{t_2}\rfloor\}$ drawn according to the Ewens measure with parameter $t_1$ and let $K(n, t)$ denote the number of its cycles, where $t\equiv (t_1, t_2)\in\mathbb [0, 1]^2$. Next,…
We investigate the $\Lambda$-Seed-Bank-Wright-Fisher process, a model describing allele frequency dynamics in populations exhibiting both skewed offspring distributions and dormancy. By performing a change of measure, we condition this…
We investigate the infinitely many demes limit of the genealogy of a sample of individuals from a subdivided population subject to sporadic mass extinction events. By exploiting a separation of timescales property of Wright's island model,…
Distances between sequences based on their $k$-mer frequency counts can be used to reconstruct phylogenies without first computing a sequence alignment. Past work has shown that effective use of k-mer methods depends on 1) model-based…
As researchers collect increasingly large molecular data sets to reconstruct the Tree of Life, the heterogeneity of signals in the genomes of diverse organisms poses challenges for traditional phylogenetic analysis. A class of phylogenetic…
In population genetics, there is often interest in inferring selection coefficients. This task becomes more challenging if multiple linked selected loci are considered simultaneously. For such a situation, we propose a novel generalized…
We consider a system of interacting Fisher-Wright diffusions with seed-bank. Individuals live in colonies and are subject to resampling and migration as long as they are active. Each colony has a structured seed-bank into which individuals…
We consider a multi-colony version of the Wright-Fisher model with seed-bank that was recently introduced by Blath et al. Individuals live in colonies and change type via resampling and mutation. Each colony contains a seed-bank that acts…
We consider the evolution of an asexually reproducing population in an uncorrelated random fitness landscape in the limit of infinite genome size, which implies that each mutation generates a new fitness value drawn from a probability…
Variation in a sample of molecular sequence data informs about the past evolutionary history of the sample's population. Traditionally, Bayesian modeling coupled with the standard coalescent, is used to infer the sample's bifurcating…
We propose a new hierarchy of semidefinite programming relaxations for inference problems. As test cases, we consider the problem of community detection in block models. The vertices are partitioned into $k$ communities, and a graph is…
The increasing availability of population-level allele frequency data across one or more related populations necessitates the development of methods that can efficiently estimate population genetics parameters, such as the strength of…