English
Related papers

Related papers: Strong Consistency of the Good-Turing Estimator

200 papers

We consider the problem of estimating the probability of an observed string drawn i.i.d. from an unknown distribution. The key feature of our study is that the length of the observed string is assumed to be of the same order as the size of…

Information Theory · Computer Science 2007-07-13 Aaron B. Wagner , Pramod Viswanath , Sanjeev R. Kulkarni

Estimating a large alphabet probability distribution from a limited number of samples is a fundamental problem in machine learning and statistics. A variety of estimation schemes have been proposed over the years, mostly inspired by the…

Machine Learning · Statistics 2018-08-20 Amichai Painsky , Meir Feder

Given $n$ samples from a population of individuals belonging to different types with unknown proportions, how do we estimate the probability of discovering a new type at the $(n+1)$-th draw? This is a classical problem in statistics,…

Statistics Theory · Mathematics 2018-06-27 Fadhel Ayed , Marco Battiston , Federico Camerlenghi , Stefano Favaro

The Good-Turing (GT) estimator for the missing mass (i.e., total probability of missing symbols) in $n$ samples is the number of symbols that appeared exactly once divided by $n$. For i.i.d. samples, the bias and squared-error risk of the…

Information Theory · Computer Science 2023-05-30 Prafulla Chandra , Andrew Thangaraj , Nived Rajaraman

Large sample size equivalence between the celebrated {\it approximated} Good-Turing estimator of the probability to discover a species already observed a certain number of times (Good, 1953) and the modern Bayesian nonparametric counterpart…

Statistics Theory · Mathematics 2019-01-29 Annalisa Cerquetti

The problem of estimating discovery probabilities originated in the context of statistical ecology, and in recent years it has become popular due to its frequent appearance in challenging applications arising in genetics, bioinformatics,…

Methodology · Statistics 2015-06-17 Stefano Favaro , Bernardo Nipoti , Yee Whye Teh

Quantifying convergence and sufficient sampling of macromolecular molecular dynamics simulations is more often than not a source of controversy (and of various ad hoc solutions) in the field. Clearly, the only reasonable, consistent and…

Quantitative Methods · Quantitative Biology 2013-12-23 Panagiotis I. Koukos , Nicholas M. Glykos

The stochastic block model (SBM) is a probabilistic model de- signed to describe heterogeneous directed and undirected graphs. In this paper, we address the asymptotic inference on SBM by use of maximum- likelihood and variational…

Statistics Theory · Mathematics 2012-10-02 Alain Celisse , J. -J. Daudin , Laurent Pierre

In probability theory, there is a tendency to treat one random variable with a given distribution as being just as good as any other. By and large this is fine because probability is (mostly) concerned with distributional properties of…

Probability · Mathematics 2013-01-31 Douglas Rizzolo

We consider optimal stopping problems, in which a sequence of independent random variables is drawn from a known continuous density. The objective of such problems is to find a procedure which maximizes the expected reward; this is often…

Probability · Mathematics 2020-12-07 Hugh Entwistle , Christopher Lustri , Georgy Sofronov

We consider the challenging problem of statistical inference for exponential-family random graph models based on a single observation of a random graph with complex dependence. To facilitate statistical inference, we consider random graphs…

Statistics Theory · Mathematics 2020-03-13 Michael Schweinberger

We consider the problem of estimating the distribution underlying an observed sample of data. Instead of maximum likelihood, which maximizes the probability of the ob served values, we propose a different estimate, the high-profile…

Artificial Intelligence · Computer Science 2012-07-19 Alon Orlitsky , Narayana Santhanam , Krishnamurthy Viswanathan , Junan Zhang

We study statistical inference and distributionally robust solution methods for stochastic optimization problems, focusing on confidence intervals for optimal values and solutions that achieve exact coverage asymptotically. We develop a…

Machine Learning · Statistics 2018-07-03 John Duchi , Peter Glynn , Hongseok Namkoong

The brilliant method due to Good and Turing allows for estimating objects not occurring in a sample. The problem, known under names "sample coverage" or "missing mass" goes back to their cryptographic work during WWII, but over years has…

Machine Learning · Statistics 2021-04-16 Maciej Skorski

In this work we study the estimation of the density of a totally positive random vector. Total positivity of the distribution of a random vector implies a strong form of positive dependence between its coordinates and, in particular, it…

Statistics Theory · Mathematics 2023-05-10 Ali Zartash , Elina Robeva

We present new results for consistency of maximum likelihood estimators with a focus on multivariate mixed models. Our theory builds on the idea of using subsets of the full data to establish consistency of estimators based on the full…

Statistics Theory · Mathematics 2019-02-13 Karl Oskar Ekvall , Galin L. Jones

In finite mixtures of location-scale distributions, if there is no constraint or penalty on the parameters, then the maximum likelihood estimator does not exist because the likelihood is unbounded. To avoid this problem, we consider a…

Statistics Theory · Mathematics 2011-03-04 Kentaro Tanaka

When faced with a small sample from a large universe of possible outcomes, scientists often turn to the venerable Good--Turing estimator. Despite its pedigree, however, this estimator comes with considerable drawbacks, such as the need to…

Statistics Theory · Mathematics 2025-09-10 Yanjun Han , Jonathan Niles-Weed , Yandi Shen , Yihong Wu

In this work, we revisit the problem of uniformity testing of discrete probability distributions. A fundamental problem in distribution testing, testing uniformity over a known domain has been addressed over a significant line of works, and…

Data Structures and Algorithms · Computer Science 2017-08-17 Tuğkan Batu , Clément L. Canonne

We consider the urn setting with two different objects, ``good'' and ``bad'', and analyze the number of draws without replacement until a good object is picked. Although the expected number of draws for this setting is a standard textbook…

Probability · Mathematics 2014-04-07 John Ahlgren
‹ Prev 1 2 3 10 Next ›