Related papers: Concentration inequalities for the sum in sampling…
Nonprobability (convenience) samples are increasingly sought to stabilize estimations for one or more population variables of interest that are performed using a randomized survey (reference) sample by increasing the effective sample size.…
Let $\mathcal{P}$ be a subset of primes and for each prime $p\in \mathcal{P}$, consider a subset $\mathcal{L}_p$ of $\mathbb{Z}/p\mathbb{Z}$. We provide restriction estimates with integers $\leq N$ sifted by…
We consider the problem of integrating a small probability sample (ps) and a non-probability sample (nps). By definition, for the nps, there are no survey weights, but for the ps, there are survey weights. The key issue is that the nps,…
In this note, we show that the relative entropy of an empirical distribution of $n$ samples drawn from a set of size $k$ with respect to the true underlying distribution is exponentially concentrated around its expectation, with central…
We provide bounds on the tail probabilities for simple procedures that generate random samples _without replacement_, when the probabilities of being selected need not be equal.
Let $\Xi_n=\{\xi_1,\dots,\xi_n\}$ be a sample of $n$ independent points distributed in a regular closed element $K$ of the extended convex ring in $\mathbb{R}^d$ according to a probability measure $\mu$ on $K$, admitting a density function.…
In this paper, we establish a new law of large numbers with the rate of convergence for special partial sums in a probability space. The proof relies on nonlinear expectation theory, as the uncertainty of random variables in the special…
We apply a common measure of randomness, the entropy, in the context of iterated functions on a finite set with n elements. For a permutation, it turns out that this entropy is asymptotically (for a growing number of iterations) close to…
Data represented by probability measures arise as empirical distributions, posterior distributions, and feature-based representations of complex objects. We study heterogeneity in a population of probability measures through the expected…
Compression of integer sets and sequences has been extensively studied for settings where elements follow a uniform probability distribution. In addition, methods exist that exploit clustering of elements in order to achieve higher…
Let $M_n^{(k)}$ denote the $k$th largest maximum of a sample $(X_1,X_2,...,X_n)$ from parent $X$ with continuous distribution. Assume there exist normalizing constants $a_n>0$, $b_n\in \mathbb{R}$ and a nondegenerate distribution $G$ such…
In "Recognizing the Maximum of a Sequence", Gilbert and Mosteller analyze a full information game where n measurements from an uniform distribution are drawn and a player (knowing n) must decide at each draw whether or not to choose that…
F\'elix-Medina and Thompson (2004) proposed a variant of link-tracing sampling to estimate the size of a hidden population such as drug users, sexual workers or homeless people. In their variant a sampling frame of sites where the members…
Corresponding to $n$ independent non-negative random variables $X_1,...,X_n$, are values $M_1,...,M_n$, where each $M_i$ is the expected value of the maximum of $n$ independent copies of $X_i$. We obtain an upper bound to the expected value…
We study the relative entropy between the empirical estimate of a discrete distribution and the true underlying distribution. If the minimum value of the probability mass function exceeds an $\alpha > 0$ (i.e. when the true underlying…
We study sums of a random multiplicative function; this is an example, of number-theoretic interest, of sums of products of independent random variables (chaoses). Using martingale methods, we establish a normal approximation for the sum…
Quantization for a probability distribution refers to the idea of estimating a given probability by a discrete probability supported by a finite number of points. In this paper, firstly a general approach to this process is outlined using…
Let P_{n,d,D} denote the graph taken uniformly at random from the set of all labelled planar graphs on {1,2,...,n} with minimum degree at least d(n) and maximum degree at most D(n). We use counting arguments to investigate the probability…
We consider the problem of stable sampling of multivariate real polynomials of large degree in a general framework where the polynomials are defined on an affine real algebraic variety $M$, equipped with a weighted measure. In particular,…
Many testing problems are readily amenable to randomised tests such as those employing data splitting. However despite their usefulness in principle, randomised tests have obvious drawbacks. Firstly, two analyses of the same dataset may…