English
Related papers

Related papers: On the Concentration of the Missing Mass

200 papers

Nonprobability (convenience) samples are increasingly sought to stabilize estimations for one or more population variables of interest that are performed using a randomized survey (reference) sample by increasing the effective sample size.…

We study $|A + A|$ as a random variable, where $A \subseteq \{0, \dots, N\}$ is a random subset such that each $0 \le n \le N$ is included with probability $0 < p < 1$, and where $A + A$ is the set of sums $a + b$ for $a,b$ in $A$. Lazarev,…

Number Theory · Mathematics 2024-02-02 Aditya Jambhale , Rauan Kaldybayev , Steven J. Miller , Chris Yao

Determining the strength of non-linear statistical dependencies between two variables is a crucial matter in many research fields. The established measure for quantifying such relations is the mutual information. However, estimating mutual…

Data Analysis, Statistics and Probability · Physics 2019-07-24 Damián G. Hernández , Inés Samengo

We overview results on the topic of Poisson approximation that are missed in existing surveys. The topic of Poisson approximation to the distribution of a sum of integer-valued random variables is presented as well. We do not restrict…

Probability · Mathematics 2019-04-12 S. Y. Novak

Given a finite set of integers $A$, its sumset is $A+A:= \{a_i+a_j \mid a_i,a_j\in A\}$. We examine $|A+A|$ as a random variable, where $A\subset I_n = [0,n-1]$, the set of integers from 0 to $n-1$, so that each element of $I_n$ is in $A$…

Information theory is built on probability measures and by definition a probability measure has total mass 1. Probability measures are used to model uncertainty, and one may ask how important it is that the total mass is one. We claim that…

Information Theory · Computer Science 2022-02-08 Peter Harremoës

In causal inference, sensitivity analysis is important to assess the robustness of study conclusions to key assumptions. We perform sensitivity analysis of the assumption that missing outcomes are missing completely at random. We follow a…

Statistics Theory · Mathematics 2023-05-12 Bart Eggen , Stéphanie L. van der Pas , Aad W. van der Vaart

This article focuses on measurement error in covariates in regression analyses in which the aim is to estimate the association between one or more covariates and an outcome, adjusting for confounding. Error in covariate measurements, if…

Methodology · Statistics 2019-10-16 Ruth H. Keogh , Jonathan W. Bartlett

Selecting between competing statistical models is a challenging problem especially when the competing models are non-nested. In this paper we offer a simple solution by devising an algorithm which combines MCMC and importance sampling to…

Power-law distributions are typical macroscopic features occurring in almost all complex systems observable in nature. As a result, researchers in quantitative analyses must often generate random synthetic variates obeying power-law…

Physics and Society · Physics 2014-11-11 Filippo Radicchi

We study high-dimensional sample covariance matrices based on independent random vectors with missing coordinates. The presence of missing observations is common in modern applications such as climate studies or gene expression…

Probability · Mathematics 2016-03-01 Kamil Jurczak , Angelika Rohde

In the setting where we have $n$ independent observations of a random variable $X$, we derive explicit error bounds in total variation distance when approximating the number of observations equal to the maximum of the sample (in the case…

Probability · Mathematics 2026-04-10 Fraser Daly

Driven by applications in telecommunication networks, we explore the simulation task of estimating rare event probabilities for tandem queues in their steady state. Existing literature has recognized that importance sampling methods can be…

Machine Learning · Computer Science 2025-04-22 Ruoning Zhao , Xinyun Chen

Selection bias is a common concern in epidemiologic studies. In the literature, selection bias is often viewed as a missing data problem. Popular approaches to adjust for bias due to missing data, such as inverse probability weighting, rely…

We consider the problem of closeness testing for two discrete distributions in the practically relevant setting of \emph{unequal} sized samples drawn from each of them. Specifically, given a target error parameter $\varepsilon > 0$, $m_1$…

Machine Learning · Computer Science 2015-04-20 Bhaswar B. Bhattacharya , Gregory Valiant

Under the null hypothesis, the marginal probability of the positive response is symmetric at any specified correlated coefficient, and the discordance probability is also symmetric to the positive response probability. The marginal…

Methodology · Statistics 2022-11-11 Guanghui Huang

Suppose $X_1,\dots, X_n$ is a random sample from a bounded and decreasing density $f_0$ on $[0,\infty)$. We are interested in estimating such $f_0$, with special interest in $f_0(0)$. This problem is encountered in various statistical…

Statistics Theory · Mathematics 2020-09-14 Geurt Jongbloed , Frank van der Meulen , Lixue Pang

A random set is a generalisation of a random variable, i.e. a set-valued random variable. The random set theory allows a unification of other uncertainty descriptions such as interval variable, mass belief function in Dempster-Shafer theory…

Numerical Analysis · Mathematics 2018-11-27 Truong-Vinh Hoang , Hermann G. Matthies

For frequentist settings in which parameter randomness represents variability rather than uncertainty, the ideal measure of the support for one hypothesis over another is the difference in the posterior and prior log odds. For situations in…

Statistics Theory · Mathematics 2013-09-03 David R. Bickel

Stein's (1959) problem highlights the phenomenon called the probability dilution in high dimensional cases, which is known as a fundamental deficiency in probabilistic inference. The satellite conjunction problem also suffers from…

Methodology · Statistics 2023-10-17 Hangbin Lee , Youngjo Lee