Related papers: Universal features of surname distribution in a su…
We consider a Moran-type model of cultural evolution, which describes how traits emerge, are transmitted, and get lost in populations. Our analysis focuses on the underlying cultural genealogies; they were first described by Aguilar and…
For a pair consisting of a gene tree and a species tree, the ancestral configurations at an internal node of the species tree are the distinct sets of gene lineages that can be present at that node. Ancestral configurations appear in…
Using demographic data of high spatial resolution for a region in the south of Europe, we study the population over fixed-size spatial cells. We find that, counterintuitively, the distribution of the number of inhabitants per cell increases…
Large whole-genome sequencing projects have provided access to much of the rare variation in human populations, which is highly informative about population structure and recent demography. Here, we show how the age of rare variants can be…
The wide availability of biological data at the genome-scale and across multiple variables has resulted in statistical questions regarding the enrichment or depletion of the number of discrete objects (e.g. genes) identified in individual…
We consider an exactly solvable model of branching random walk with random selection, which describes the evolution of a population with $N$ individuals on the real line. At each time step, every individual reproduces independently, and its…
Perfect sorting by reversals, a problem originating in computational genomics, is the process of sorting a signed permutation to either the identity or to the reversed identity permutation, by a sequence of reversals that do not break any…
We analyze several florae (collections of plant species populating specific areas) in different geographic and climatic regions. For every list of species we produce a taxonomic classification tree and we consider its statistical…
Various approaches to alignment-free sequence comparison are based on the length of exact or inexact word matches between two input sequences. Haubold {\em et al.} (2009) showed how the average number of substitutions between two DNA…
Shannon information (SI) and its special case, divergence, are defined for a DNA sequence in terms of probabilities of chemical words in the sequence and are computed for a set of complete genomes highly diverse in length and composition.…
Commonly observed patterns typically follow a few distinct families of probability distributions. Over one hundred years ago, Karl Pearson provided a systematic derivation and classification of the common continuous distributions. His…
Many questions that we have about the history and dynamics of organisms have a geographical component: How many are there, and where do they live? How do they move and interbreed across the landscape? How were they moving a thousand years…
Starting from a master equation, we derive the evolution equation for the size distribution of elements in an evolving system, where each element can grow, divide into two, and produce new elements. We then probe general solutions of the…
Genotyping errors are known to influence the power of both family-based and case-control studies in the genetics of complex disease. Estimating genotyping error rate in a given dataset can be complex, but when family information is…
We study the probabilistic evolution of a birth and death continuous time measure-valued process with mutations and ecological interactions. The individuals are characterized by (phenotypic) traits that take values in a compact metric…
We investigate stochastic comparisons between exponential family distributions and their mixtures with respect to the usual stochastic order, the hazard rate order, the reversed hazard rate order, and the likelihood ratio order. A general…
The evolutionary dynamics of molecular populations are strongly dependent on the structure of genotype spaces. The map between genotype and phenotype determines how easily genotype spaces can be navigated and the accessibility of…
We examine the extent to which random samplings from the values of a random set, determine the distribution of the random set itself. We also comment on how, given the statistics of the sampling, to detect the distribution. Several methods…
We investigate two stochastic models of a growing population subject to selection and mutation. In our models each individual carries a fitness which determines its mean offspring number. Many of these offspring inherit their parent's…
The Law of Large Numbers tells us that as the sample size (N) is increased, the sample mean converges on the population mean, provided that the latter exists. In this paper, we investigate the opposite effect: keeping the sample size fixed…