Related papers: Concentration inequalities for the sum in sampling…
Let X_n=(x_{ij}) be an n by p data matrix, where the n rows form a random sample of size n from a certain p-dimensional population distribution. Let R_n=(\rho_{ij}) be the p\times p sample correlation matrix of X_n; that is, the entry…
Maximum entropy method is a constructive criterion for setting up a probability distribution maximally non-committal to missing information on the basis of partial knowledge, usually stated as constrains on expectation values of some…
In this paper we describe the alternative approach to the sample boundedness and continuity of stochastic processes. We show that the regularity of paths can be understood in terms of a distribution of the argument maximum. For a centered…
The Shannon entropy, and related quantities such as mutual information, can be used to quantify uncertainty and relevance. However, in practice, it can be difficult to compute these quantities for arbitrary probability distributions,…
This paper presents theoretical results on combining non-probability and probability survey samples through mass imputation, an approach originally proposed by Rivers (2007) as sample matching without rigorous theoretical justification.…
In many empirical studies of a large two-sided matching market (such as in a college admissions problem), the researcher performs statistical inference under the assumption that they observe a random sample from a large matching market. In…
In classical statistics and distribution testing, it is often assumed that elements can be sampled from some distribution $P$, and that when an element $x$ is sampled, the probability $P$ of sampling $x$ is also known. Recent work in…
We investigate quantitative implications of the notion of log-concavity through a probabilistic interpretation. In particular, we derive concentration inequalities, moment and entropy bounds for random variables satisfying a precise degree…
This paper studies majorization of high tensor powers of finitely supported probability distributions. Viewing probability distributions as a resource with majorization as a means of transformation corresponds to the resource theory of pure…
The basic goal of quantization for probability distribution is to reduce the number of values, which is typically uncountable, describing a probability distribution to some finite set and thus approximation of a continuous probability…
The main idea of nested sampling is to substitute the high-dimensional likelihood integral over the parameter space $\Omega$ by an integral over the unit line $[0,1]$ by employing a push-forward with respect to a suitable transformation.…
Information theory is built on probability measures and by definition a probability measure has total mass 1. Probability measures are used to model uncertainty, and one may ask how important it is that the total mass is one. We claim that…
We primarily investigate congruences modulo $p$ for finite sums of the form $\sum_k\binom{rk}{k}x^k/k$ over the ranges $0<k<p$ and $0<k<p/r$, where $p$ is a prime larger than the positive integer $r$. Here $x$ is an indeterminate, thus…
Let $N_n=\{1,2,...,n\}$. Elements are drawn from the set $N_n$ with replacement, assuming that each element has probability $1/n$ of being drawn. We determine the limiting distributions for the waiting time until the given portion of pairs…
This paper is devoted to establishing exponential bounds for the probabilities of deviation of a sample sum from its expectation, when the variables involved in the summation are obtained by sampling in a finite population according to a…
We study expected values of the polynomials $P_N^{}(z)=\prod_{1\leq n\leq N}(X_n^2+z^2)$ whose $2N$ zeros $\{\pm i X_k\}^{}_{k=1,...,N}$ are generated by $N$ identically distributed multi-variate mean-zero normal random variables…
Some practical results are derived for population inference based on a sample, under the two qualitative conditions of 'ignorability' and exchangeability. These are the 'Histogram Theorem', for predicting the outcome of a non-sampled member…
We study the use of sampling for efficiently mining the top-K frequent itemsets of cardinality at most w. To this purpose, we define an approximation to the top-K frequent itemsets to be a family of itemsets which includes (resp., excludes)…
We consider the sum of power weighted nearest neighbor distances in a sample of size n from a multivariate density f of possibly unbounded support. We give various criteria guaranteeing that this sum satisfies a law of large numbers for…
We study asymptotic probabilities of attaining the maximum in heterogeneous Gaussian samples. In the two-group setting, the first sample has variance $1$ and size $n_1$, while the second has variance $\sigma^2>1$ and size $n_2$. We…