English
Related papers

Related papers: Missing $g$-mass: Investigating the Missing Parts …

200 papers

The missing mass refers to the probability of elements not observed in a sample, and since the work of Good and Turing during WWII, has been studied extensively in many areas including ecology, linguistic, networks and information theory.…

Information Theory · Computer Science 2021-04-16 Maciej Skorski

Distribution estimation under error-prone or non-ideal sampling modelled as "sticky" channels have been studied recently motivated by applications such as DNA computing. Missing mass, the sum of probabilities of missing letters, is an…

Statistics Theory · Mathematics 2022-02-08 Prafulla Chandra , Andrew Thangaraj , Nived Rajaraman

Given $n$ samples from a population of individuals belonging to different types with unknown proportions, how do we estimate the probability of discovering a new type at the $(n+1)$-th draw? This is a classical problem in statistics,…

Statistics Theory · Mathematics 2018-06-27 Fadhel Ayed , Marco Battiston , Federico Camerlenghi , Stefano Favaro

Consider a random sample $(X_{1},\ldots,X_{n})$ from an unknown discrete distribution $P=\sum_{j\geq1}p_{j}\delta_{s_{j}}$ on a countable alphabet $\mathbb{S}$, and let $(Y_{n,j})_{j\geq1}$ be the empirical frequencies of distinct symbols…

Statistics Theory · Mathematics 2024-07-12 Stefano Favaro , Zacharie Naulet

The Good-Turing (GT) estimator for the missing mass (i.e., total probability of missing symbols) in $n$ samples is the number of symbols that appeared exactly once divided by $n$. For i.i.d. samples, the bias and squared-error risk of the…

Information Theory · Computer Science 2023-05-30 Prafulla Chandra , Andrew Thangaraj , Nived Rajaraman

We study the estimation and concentration on its expectation of the probability to observe data further than a specified distance from a given iid sample in a metric space. The problem extends the classical problem of estimation of the…

Statistics Theory · Mathematics 2022-11-23 Andreas Maurer

The missing mass refers to the proportion of data points in an unknown population of classifier inputs that belong to classes not present in the classifier's training data, which is assumed to be a random sample from that unknown…

Machine Learning · Computer Science 2025-03-11 Seongmin Lee , Marcel Böhme

The brilliant method due to Good and Turing allows for estimating objects not occurring in a sample. The problem, known under names "sample coverage" or "missing mass" goes back to their cryptographic work during WWII, but over years has…

Machine Learning · Statistics 2021-04-16 Maciej Skorski

We consider the problem of estimating the missing mass, partition function or evidence and its probability distribution in the case that for each sample point in the discrete sample space its (unnormalized) probability mass is revealed.…

Statistics Theory · Mathematics 2026-03-16 Bastiaan J. Braams

This paper gives a review of concentration inequalities which are widely employed in non-asymptotical analyses of mathematical statistics in a wide range of settings, from distribution-free to distribution-dependent, from sub-Gaussian to…

Statistics Theory · Mathematics 2025-02-24 Huiming Zhang , Song Xi Chen

The problem of estimating the missing mass or total probability of unseen elements in a sequence of $n$ random samples is considered under the squared error loss function. The worst-case risk of the popular Good-Turing estimator is shown to…

Information Theory · Computer Science 2017-05-16 Nikhilesh Rajaraman , Andrew Thangaraj , Ananda Theertha Suresh

Novel concentration inequalities are obtained for the missing mass, i.e. the total probability mass of the outcomes not observed in the sample. We derive distribution-free deviation bounds with sublinear exponents in deviation size for…

Machine Learning · Statistics 2015-12-02 Bahman Yari Saeed Khanloo

A random variable is sampled from a discrete distribution. The missing mass is the probability of the set of points not observed in the sample. We sharpen and simplify McAllester and Ortiz's results (JMLR, 2003) bounding the probability of…

Probability · Mathematics 2012-10-12 Daniel Berend , Aryeh Kontorovich

We study the problem of estimating a distribution over a finite alphabet from an i.i.d. sample, with accuracy measured in relative entropy (Kullback-Leibler divergence). While optimal bounds on the expected risk are known, high-probability…

Statistics Theory · Mathematics 2026-02-27 Jaouad Mourtada

The generalized negative binomial distribution (GNB) is a new flexible family of discrete distributions that are mixed Poisson laws with the mixing generalized gamma (GG) distributions. This family of discrete distributions is very wide and…

Methodology · Statistics 2018-09-18 Andrey K. Gorshenin , Victor Yu. Korolev

The problem of missing mass in statistical inference (posed by McAllester and Ortiz, NIPS'02; most recently revisited by Changa and Thangaraj, ISIT'2019) seeks to estimate the weight of symbols that have not been sampled yet from a source.…

Probability · Mathematics 2020-01-15 Maciej Skorski

We give tight lower and upper bounds on the expected missing mass for distributions over finite and countably infinite spaces. An essential characterization of the extremal distributions is given. We also provide an extension to totally…

Statistics Theory · Mathematics 2011-11-10 Daniel Berend , Aryeh Kontorovich

Consider a pair of cumulative distribution functions $F$ and $G$, where $F$ is unknown and $G$ is a known reference distribution. Given a sample from $F$, we propose tests to detect the convexity or the concavity of $G^{-1}\circ F$ versus…

Statistics Theory · Mathematics 2025-06-25 Tommaso Lando , Mohammed Es-Salih Benjrada

We consider the problem of estimating the probability of an observed string drawn i.i.d. from an unknown distribution. The key feature of our study is that the length of the observed string is assumed to be of the same order as the size of…

Information Theory · Computer Science 2007-07-13 Aaron B. Wagner , Pramod Viswanath , Sanjeev R. Kulkarni

We revisit the problem of estimating the mean of a real-valued distribution, presenting a novel estimator with sub-Gaussian convergence: intuitively, "our estimator, on any distribution, is as accurate as the sample mean is for the Gaussian…

Statistics Theory · Mathematics 2020-11-18 Jasper C. H. Lee , Paul Valiant
‹ Prev 1 2 3 10 Next ›