Related papers: Exact formulas for the variance of several balance…
We present a general framework for balancing expressions (terms) in form of so called tree straight-line programs. The latter can be seen as circuits over the free term algebra extended by contexts (terms with a hole) and the operations…
We consider the problem of identifying jointly the ancestral sequence, the phylogeny and the parameters in models of DNA sequence evolution with insertion and deletion (indel). Under the classical TKF91 model of sequence evolution, we…
We introduce a model for the evolution of species triggered by generation of novel features and exhaustive combination with other available traits. Under the assumption that innovations are rare, we obtain a bursty branching process of…
Because biological processes can make different loci have different evolutionary histories, species tree estimation requires multiple loci from across the genome. While many processes can result in discord between gene trees and species…
We introduce a scale-free method for testing the proportionality of branch lengths between two phylogenetic trees that have the same topology and contain the same set of taxa. This method scales both trees to a total length of 1 and sums up…
Covariate balance is crucial for unconfounded descriptive or causal comparisons. However, lack of balance is common in observational studies. This article considers weighting strategies for balancing covariates. We define a general class of…
Covariate balance is crucial for unconfounded descriptive or causal comparisons. However, lack of balance is common in observational studies. This article considers weighting strategies for balancing covariates. We define a general class of…
We give a short proof of Cayley's tree formula for counting the number of different labeled trees on $n$ vertices. The following nonlinear recursive relation for the number of labeled trees on $n$ vertices is deduced from a combinatorial…
This paper deals with statistical inference for the scale mixture models. We study an estimation approach based on the Mellin -- Stieltjes transform that can be applied to both discrete and absolute continuous mixing distributions. The…
We study numerical invariants associated with the reduction of singularities of holomorphic foliation germs on $(\mathbb{C}^2, 0)$. Building on our previous work on generalized curve foliations, we extend explicit formulas for several…
Comparing alternatives in pairs is a very well known technique of ranking creation. The answer to how reliable and trustworthy ranking is depends on the inconsistency of the data from which it was created. There are many indices used for…
A Gaussian fluctuation formula is proved for linear statistics of complex random matrices in the case that the statistic is rotationally invariant. For a general linear statistic without this symmetry, Coulomb gas theory is used to predict…
Phylogenetic diversity indices provide a formal way to apportion 'evolutionary heritage' across species. Two natural diversity indices are Fair Proportion (FP) and Equal Splits (ES). FP is also called 'evolutionary distinctiveness' and, for…
In this paper, I proof that Importance Sampling estimates based on dependent sample sets are consistent under certain conditions. This can be used to reduce variance in Bayesian Models with factorizing likelihoods, using sample sets that…
Yule's 1925 paper introducing the branching model that bears his name was a landmark contribution to the biodiversity sciences. In his paper, Yule developed stochastic models to explain the observed distribution of species across genera and…
An occupancy problem with an infinite number of bins and a random probability vector for the locations of the balls is considered. The respective sizes of bins are related to the split times of a Yule process. The asymptotic behavior of the…
If predictions for species extinctions hold, then the `tree of life' today may be quite different to that in (say) 100 years. We describe a technique to quantify how much each species is likely to contribute to future biodiversity, as…
We investigate saddlepoint approximations applied to the score test statistic in genome-wide association studies with binary phenotypes. The inaccuracy in the normal approximation of the score test statistic increases with increasing sample…
On trees of fixed order, we show a direct relation between Kemeny's constant and Wiener index, and provide a new formula of Kemeny's constant from the relation with a combinatorial interpretation. Moreover, the relation simplifies proofs of…
We study the statistics of height and balanced height in the binary search tree problem in computer science. The search tree problem is first mapped to a fragmentation problem which is then further mapped to a modified directed polymer…