Related papers: Overdispersed and Markovian Children
The Newcomb-Benford law, also known as the first-digit law, gives the probability distribution associated with the first digit of a dataset, so that, for example, the first significant digit has a probability of $30.1$ % of being $1$ and…
Estimating the unknown number of classes in a population has numerous important applications. In a Poisson mixture model, the problem is reduced to estimating the odds that a class is undetected in a sample. The discontinuity of the odds…
Clinicians and scientists have traditionally focussed on whether their findings will be replicated and are very familiar with the concept. The probability that a replication study yields an effect with the same sign, or the same statistical…
Faced with a sequence of N binary events, such as coin flips (or Ising spins), it is natural to ask whether these events reflect some underlying dynamic signals or are just random. Plausible models for the dynamics of hidden biases lead to…
The standard textbook method for estimating the probability of a biased coin from finite tosses implicitly assumes the sample sizes are large and gives incorrect results for small samples. We describe the exact solution, which is correct…
In this paper, we focus on stochastic comparisons of extreme order statistics stemming from multiple-outlier scale models with dependence. Archimedean copula is used to model dependence structure among nonnegative random variables.…
Evolutionary theory predicts that children fare better if they resemble their father. However, if a man is promiscuous then his children tend to be (unwittingly) raised within other families; such children should fare better if they do not…
We adopt a physically motivated empirical approach to the characterisation of the distributions of twin and triplet primes within the set of primes, rather than in the set of all natural numbers. Remarkably, the occurrences of twins or…
We approximate the distribution of the sum of independent but not necessarily identically distributed Bernoulli random variables using a shifted binomial distribution where the three parameters (the number of trials, the probability of…
We investigate statistical inference across time scales. We take as toy model the estimation of the intensity of a discretely observed compound Poisson process with symmetric Bernoulli jumps. We have data at different time scales:…
The sample frequency spectrum (SFS) of DNA sequences from a collection of individuals is a summary statistic which is commonly used for parametric inference in population genetics. Despite the popularity of SFS-based inference methods,…
A statistical generalization is made of microeconomics in the spirit of going from classical to statistical mechanics. The price and quantity of every commodity1 traded in the market, at each instant of time, is considered to be an…
Throughout history most young adults have chosen to live where their parents did while a smaller number moved away. This is sufficient, by proof and simulation, to account for the well-known power law distributions of city sizes. The model…
A classical problem in statistics is estimating the expected coverage of a sample, which has had applications in gene expression, microbial ecology, optimization, and even numismatics. Here we consider a related extension of this problem to…
A common approach to statistical learning with big-data is to randomly split it among $m$ machines and learn the parameter of interest by averaging the $m$ individual estimates. In this paper, focusing on empirical risk minimization, or…
When humans infer underlying probabilities from stochastic observations, they exhibit biases and variability that cannot be explained on the basis of sound, Bayesian manipulations of probability. This is especially salient when beliefs are…
We propose some new results on the comparison of the minimum or maximum order statistic from a random number of non-identical random variables. Under the non-identical set-up, with certain conditions, we prove that random minimum (maximum)…
We propose a new method for the calculation of the statistical properties, as e.g. the entropy, of unknown generators of symbolic sequences. The probability distribution p(k) of the elements k of a population can be approximated by the…
Graded posets frequently arise throughout combinatorics, where it is natural to try to count the number of elements of a fixed rank. These counting problems are often $\#\textbf{P}$-complete, so we consider approximation algorithms for…
Dispersal is ubiquitous throughout the tree of life: factors selecting for dispersal include kin competition, inbreeding avoidance and spatiotemporal variation in resources or habitat suitability. These factors differ in whether they…