Related papers: Ordered and size-biased frequencies in GEM and Gib…
Let $X_1,\ldots,X_n$ be a random sample from the Gamma distribution with density $f(x)=\lambda^{\alpha}x^{\alpha-1}e^{-\lambda x}/\Gamma(\alpha)$, $x>0$, where both $\alpha>0$ (the shape parameter) and $\lambda>0$ (the reciprocal scale…
Order statistics of periodic, Gaussian noise with 1/f^{\alpha} power spectrum is investigated. Using simulations and phenomenological arguments, we find three scaling regimes for the average gap d_k=<x_k-x_{k+1}> between the k-th and…
We present a number models describing the sequential deposition of a mixture of particles whose size distribution is determined by the power-law $p(x) \sim \alpha x^{\alpha-1}$, $x\leq l$ . We explicitly obtain the scaling function in the…
The gapped local alignment score of two random sequences follows a Gumbel distribution. If computers could estimate the parameters of the Gumbel distribution within one second, the use of arbitrary alignment scoring schemes could increase…
We provide an information-theoretic analysis of the generalization ability of Gibbs-based transfer learning algorithms by focusing on two popular transfer learning approaches, $\alpha$-weighted-ERM and two-stage-ERM. Our key result is an…
We consider the problem of learning a target probability distribution over a set of $N$ binary variables from the knowledge of the expectation values (with this target distribution) of $M$ observables, drawn uniformly at random. The space…
The order $O_n(\sigma)$ of a permutation $\sigma$ of $n$ objects is the smallest integer $k \geq 1$ such that the $k$-th iterate of $\sigma$ gives the identity. A remarkable result about the order of a uniformly chosen permutation is due to…
Gibbs-type random probability measures and the exchangeable random partitions they induce represent the subject of a rich and active literature. They provide a probabilistic framework for a wide range of theoretical and applied problems…
We consider the fragmentation process with mass loss and discuss self-similar properties of the arising structure both in time and space, focusing on dimensional analysis. This exhibits a spectrum of mass exponents $\theta$, whose exact…
The Generalized Mallows Model (GMM) is a well known family of models for ranking data. A GMM is a distribution over $\mathbb{S}_n$, the set of permutations of n objects, characterized by a location parameter $\sigma \in \mathbb{S}_n$, known…
Gibbs-ERM learning is a natural idealized model of learning with stochastic optimization algorithms (such as Stochastic Gradient Langevin Dynamics and ---to some extent--- Stochastic Gradient Descent), while it also arises in other…
Consider two batches of independent or interdependent exponentiated location-scale distributed heterogeneous random variables. This article investigates ordering results for the second-order statistics from these batches when a vector of…
Prime numbers seem to distribute among the natural numbers with no other law than that of chance, however its global distribution presents a quite remarkable smoothness. Such interplay between randomness and regularity has motivated sci-…
The mathematical properties of a family of generalized beta distribution, including beta-normal, skewed-t, log-F, beta-exponential, beta-Weibull distributions have recently been studied in several publications. This paper applies these…
Motivated by the fundamental problem of measuring species diversity, this paper introduces the concept of a cluster structure to define an exchangeable cluster probability function that governs the joint distribution of a random count and…
This paper investigates the effects of data size and frequency range on distributional semantic models. We compare the performance of a number of representative models for several test settings over data of varying sizes, and over test…
We study how the order of N independent random walks in one dimension evolves with time. Our focus is statistical properties of the inversion number m, defined as the number of pairs that are out of sort with respect to the initial…
A learned generative model often produces biased statistics relative to the underlying data distribution. A standard technique to correct this bias is importance sampling, where samples from the model are weighted by the likelihood ratio…
We derive a large deviation principle for random permutations induced by probability measures of the unit square, called permutons. These permutations are called $\mu$-random permutations. We also introduce and study a new general class of…
This paper introduces a convenient strategy for coding and predicting sequences of independent, identically distributed random variables generated from a large alphabet of size $m$. In particular, the size of the sample is allowed to be…