English
Related papers

Related papers: On the Concentration of the Missing Mass

200 papers

In this paper, we are concerned with obtaining distribution-free concentration inequalities for mixture of independent Bernoulli variables that incorporate a notion of variance. Missing mass is the total probability mass associated to the…

Machine Learning · Statistics 2015-03-05 Bahman Yari Saeed Khanloo

The missing mass refers to the probability of elements not observed in a sample, and since the work of Good and Turing during WWII, has been studied extensively in many areas including ecology, linguistic, networks and information theory.…

Information Theory · Computer Science 2021-04-16 Maciej Skorski

We study the estimation and concentration on its expectation of the probability to observe data further than a specified distance from a given iid sample in a metric space. The problem extends the classical problem of estimation of the…

Statistics Theory · Mathematics 2022-11-23 Andreas Maurer

We are concerned with obtaining novel concentration inequalities for the missing mass, i.e. the total probability mass of the outcomes not observed in the sample. We not only derive - for the first time - distribution-free Bernstein-like…

Machine Learning · Statistics 2015-06-22 Bahman Yari Saeed Khanloo , Gholamreza Haffari

Novel concentration inequalities are obtained for the missing mass, i.e. the total probability mass of the outcomes not observed in the sample. We derive distribution-free deviation bounds with sublinear exponents in deviation size for…

Machine Learning · Statistics 2015-12-02 Bahman Yari Saeed Khanloo

The missing mass refers to the proportion of data points in an unknown population of classifier inputs that belong to classes not present in the classifier's training data, which is assumed to be a random sample from that unknown…

Machine Learning · Computer Science 2025-03-11 Seongmin Lee , Marcel Böhme

This paper shows that one cannot learn the probability of rare events without imposing further structural assumptions. The event of interest is that of obtaining an outcome outside the coverage of an i.i.d. sample from a discrete…

Machine Learning · Statistics 2015-03-13 Elchanan Mossel , Mesrob I. Ohannessian

We consider the problem of estimating the missing mass, partition function or evidence and its probability distribution in the case that for each sample point in the discrete sample space its (unnormalized) probability mass is revealed.…

Statistics Theory · Mathematics 2026-03-16 Bastiaan J. Braams

Distribution estimation under error-prone or non-ideal sampling modelled as "sticky" channels have been studied recently motivated by applications such as DNA computing. Missing mass, the sum of probabilities of missing letters, is an…

Statistics Theory · Mathematics 2022-02-08 Prafulla Chandra , Andrew Thangaraj , Nived Rajaraman

Given $n$ samples from a population of individuals belonging to different types with unknown proportions, how do we estimate the probability of discovering a new type at the $(n+1)$-th draw? This is a classical problem in statistics,…

Statistics Theory · Mathematics 2018-06-27 Fadhel Ayed , Marco Battiston , Federico Camerlenghi , Stefano Favaro

The prior distribution on parameters of a sampling distribution is the usual starting point for Bayesian uncertainty quantification. In this paper, we present a different perspective which focuses on missing observations as the source of…

Methodology · Statistics 2021-11-23 Edwin Fong , Chris Holmes , Stephen G. Walker

Consider a random sample $(X_{1},\ldots,X_{n})$ from an unknown discrete distribution $P=\sum_{j\geq1}p_{j}\delta_{s_{j}}$ on a countable alphabet $\mathbb{S}$, and let $(Y_{n,j})_{j\geq1}$ be the empirical frequencies of distinct symbols…

Statistics Theory · Mathematics 2024-07-12 Stefano Favaro , Zacharie Naulet

We give tight lower and upper bounds on the expected missing mass for distributions over finite and countably infinite spaces. An essential characterization of the extremal distributions is given. We also provide an extension to totally…

Statistics Theory · Mathematics 2011-11-10 Daniel Berend , Aryeh Kontorovich

Estimating the underlying distribution from \textit{iid} samples is a classical and important problem in statistics. When the alphabet size is large compared to number of samples, a portion of the distribution is highly likely to be…

Statistics Theory · Mathematics 2023-05-30 Prafulla Chandra , Andrew Thangaraj

The problem of missing mass in statistical inference (posed by McAllester and Ortiz, NIPS'02; most recently revisited by Changa and Thangaraj, ISIT'2019) seeks to estimate the weight of symbols that have not been sampled yet from a source.…

Probability · Mathematics 2020-01-15 Maciej Skorski

The problem of estimating the missing mass or total probability of unseen elements in a sequence of $n$ random samples is considered under the squared error loss function. The worst-case risk of the popular Good-Turing estimator is shown to…

Information Theory · Computer Science 2017-05-16 Nikhilesh Rajaraman , Andrew Thangaraj , Ananda Theertha Suresh

An infinite urn scheme is defined by a probability mass function $(p_j)_{j\geq1}$ over positive integers. A random allocation consists of a sample of $N$ independent drawings according to this probability distribution where $N$ may be…

Statistics Theory · Mathematics 2016-09-29 Anna Ben-Hamou , Stéphane Boucheron , Mesrob I. Ohannessian

The interplay between missing data and model uncertainty -- two classic statistical problems -- leads to primary questions that we formally address from an objective Bayesian perspective. For the general regression problem, we discuss the…

Existing approaches to model uncertainty typically either compare models using a quantitative model selection criterion or evaluate posterior model probabilities having set a prior. In this paper, we propose an alternative strategy which…

Methodology · Statistics 2025-03-26 Vik Shirvaikar , Stephen G. Walker , Chris Holmes

Feature models are popular in machine learning and they have been recently used to solve many unsupervised learning problems. In these models every observation is endowed with a finite set of features, usually selected from an infinite…

Statistics Theory · Mathematics 2019-02-28 Fadhel Ayed , Marco Battiston , Federico Camerlenghi , Stefano Favaro
‹ Prev 1 2 3 10 Next ›