English
Related papers

Related papers: Mean-Squared Accuracy of Good-Turing Estimator

200 papers

The missing mass refers to the probability of elements not observed in a sample, and since the work of Good and Turing during WWII, has been studied extensively in many areas including ecology, linguistic, networks and information theory.…

Information Theory · Computer Science 2021-04-16 Maciej Skorski

The problem of estimating the missing mass or total probability of unseen elements in a sequence of $n$ random samples is considered under the squared error loss function. The worst-case risk of the popular Good-Turing estimator is shown to…

Information Theory · Computer Science 2017-05-16 Nikhilesh Rajaraman , Andrew Thangaraj , Ananda Theertha Suresh

Given $n$ samples from a population of individuals belonging to different types with unknown proportions, how do we estimate the probability of discovering a new type at the $(n+1)$-th draw? This is a classical problem in statistics,…

Statistics Theory · Mathematics 2018-06-27 Fadhel Ayed , Marco Battiston , Federico Camerlenghi , Stefano Favaro

Feature allocation models generalize species sampling models by allowing every observation to belong to more than one species, now called features. Under the popular Bernoulli product model for feature allocation, given $n$ samples, we…

Statistics Theory · Mathematics 2020-09-22 Fadhel Ayed , Marco Battiston , Federico Camerlenghi , Stefano Favaro

The Good-Turing (GT) estimator for the missing mass (i.e., total probability of missing symbols) in $n$ samples is the number of symbols that appeared exactly once divided by $n$. For i.i.d. samples, the bias and squared-error risk of the…

Information Theory · Computer Science 2023-05-30 Prafulla Chandra , Andrew Thangaraj , Nived Rajaraman

The missing mass refers to the proportion of data points in an unknown population of classifier inputs that belong to classes not present in the classifier's training data, which is assumed to be a random sample from that unknown…

Machine Learning · Computer Science 2025-03-11 Seongmin Lee , Marcel Böhme

Estimating a large alphabet probability distribution from a limited number of samples is a fundamental problem in machine learning and statistics. A variety of estimation schemes have been proposed over the years, mostly inspired by the…

Machine Learning · Statistics 2018-08-20 Amichai Painsky , Meir Feder

We consider the problem of estimating the probability of an observed string drawn i.i.d. from an unknown distribution. The key feature of our study is that the length of the observed string is assumed to be of the same order as the size of…

Information Theory · Computer Science 2007-07-13 Aaron B. Wagner , Pramod Viswanath , Sanjeev R. Kulkarni

Distribution estimation under error-prone or non-ideal sampling modelled as "sticky" channels have been studied recently motivated by applications such as DNA computing. Missing mass, the sum of probabilities of missing letters, is an…

Statistics Theory · Mathematics 2022-02-08 Prafulla Chandra , Andrew Thangaraj , Nived Rajaraman

Estimating the underlying distribution from \textit{iid} samples is a classical and important problem in statistics. When the alphabet size is large compared to number of samples, a portion of the distribution is highly likely to be…

Statistics Theory · Mathematics 2023-05-30 Prafulla Chandra , Andrew Thangaraj

We study the estimation and concentration on its expectation of the probability to observe data further than a specified distance from a given iid sample in a metric space. The problem extends the classical problem of estimation of the…

Statistics Theory · Mathematics 2022-11-23 Andreas Maurer

Large sample size equivalence between the celebrated {\it approximated} Good-Turing estimator of the probability to discover a species already observed a certain number of times (Good, 1953) and the modern Bayesian nonparametric counterpart…

Statistics Theory · Mathematics 2019-01-29 Annalisa Cerquetti

We consider the problem of estimating the total probability of all symbols that appear with a given frequency in a string of i.i.d. random variables with unknown distribution. We focus on the regime in which the block length is large yet no…

Information Theory · Computer Science 2016-11-15 Aaron B. Wagner , Pramod Viswanath , Sanjeev R. Kulkarni

The problem of estimating discovery probabilities originated in the context of statistical ecology, and in recent years it has become popular due to its frequent appearance in challenging applications arising in genetics, bioinformatics,…

Methodology · Statistics 2015-06-17 Stefano Favaro , Bernardo Nipoti , Yee Whye Teh

Turing's estimator allows one to estimate the probabilities of outcomes that either do not appear or only rarely appear in a given random sample. We perform a simulation study to understand the finite sample performance of several related…

Statistics Theory · Mathematics 2025-03-19 Jie Chang , Michael Grabchak , Jialin Zhang

When faced with a small sample from a large universe of possible outcomes, scientists often turn to the venerable Good--Turing estimator. Despite its pedigree, however, this estimator comes with considerable drawbacks, such as the need to…

Statistics Theory · Mathematics 2025-09-10 Yanjun Han , Jonathan Niles-Weed , Yandi Shen , Yihong Wu

A prime goal of quantum tomography is to provide quantitatively rigorous characterisation of quantum systems, be they states, processes or measurements, particularly for the purposes of trouble-shooting and benchmarking experiments in…

Quantum Physics · Physics 2015-06-12 Nathan K. Langford

We consider an original problem that arises from the issue of security analysis of a power system and that we name optimal discovery with probabilistic expert advice. We address it with an algorithm based on the optimistic paradigm and the…

Optimization and Control · Mathematics 2011-10-26 Sébastien Bubeck , Damien Ernst , Aurélien Garivier

We study the problem of estimating the parameters of a Gaussian distribution when samples are only shown if they fall in some (unknown) subset $S \subseteq \R^d$. This core problem in truncated statistics has long history going back to…

Statistics Theory · Mathematics 2019-08-06 Vasilis Kontonis , Christos Tzamos , Manolis Zampetakis

In this article we have suggested an improved estimator for estimating the population mean in simple random sampling using auxiliary information under the presence of measurement errors. The mean square error (MSE) of the proposed estimator…

Applications · Statistics 2013-12-05 Sachin Malik , Jayant Singh , Rajesh Singh
‹ Prev 1 2 3 10 Next ›