English
Related papers

Related papers: Optimal estimation of high-order missing masses, a…

200 papers

Feature models are popular in machine learning and they have been recently used to solve many unsupervised learning problems. In these models every observation is endowed with a finite set of features, usually selected from an infinite…

Statistics Theory · Mathematics 2019-02-28 Fadhel Ayed , Marco Battiston , Federico Camerlenghi , Stefano Favaro

Estimating the underlying distribution from \textit{iid} samples is a classical and important problem in statistics. When the alphabet size is large compared to number of samples, a portion of the distribution is highly likely to be…

Statistics Theory · Mathematics 2023-05-30 Prafulla Chandra , Andrew Thangaraj

We consider the problem of estimating the missing mass, partition function or evidence and its probability distribution in the case that for each sample point in the discrete sample space its (unnormalized) probability mass is revealed.…

Statistics Theory · Mathematics 2026-03-16 Bastiaan J. Braams

Distribution estimation under error-prone or non-ideal sampling modelled as "sticky" channels have been studied recently motivated by applications such as DNA computing. Missing mass, the sum of probabilities of missing letters, is an…

Statistics Theory · Mathematics 2022-02-08 Prafulla Chandra , Andrew Thangaraj , Nived Rajaraman

We study the problem of estimating the joint probability mass function (pmf) over two random variables. In particular, the estimation is based on the observation of $m$ samples containing both variables and $n$ samples missing one fixed…

Statistics Theory · Mathematics 2023-05-17 H. S. Melihcan Erol , Erixhen Sula , Lizhong Zheng

We propose a general methodology for the construction and analysis of minimax estimators for a wide class of functionals of finite dimensional parameters, and elaborate on the case of discrete distributions, where the alphabet size $S$ is…

Information Theory · Computer Science 2015-03-11 Jiantao Jiao , Kartik Venkat , Yanjun Han , Tsachy Weissman

Given $n$ samples from a population of individuals belonging to different types with unknown proportions, how do we estimate the probability of discovering a new type at the $(n+1)$-th draw? This is a classical problem in statistics,…

Statistics Theory · Mathematics 2018-06-27 Fadhel Ayed , Marco Battiston , Federico Camerlenghi , Stefano Favaro

Given $n$ samples from a population of individuals belonging to different species, what is the number $U$ of hitherto unseen species that would be observed if $\lambda n$ new samples were collected? This is an important problem in many…

Statistics Theory · Mathematics 2022-03-17 Stefano Favaro , Zacharie Naulet

We study the problem of estimating the joint probability mass function (pmf) over two random variables. In particular, the estimation is based on the observation of $m$ samples containing both variables and $n$ samples missing one fixed…

Statistics Theory · Mathematics 2024-05-16 Hasan Sabri Melihcan Erol , Lizhong Zheng

This paper addresses a problem of estimating an additive functional given $n$ i.i.d. samples drawn from a discrete distribution $P=(p_1,...,p_k)$ with alphabet size $k$. The additive functional is defined as…

Information Theory · Computer Science 2018-12-04 Kazuto Fukuchi , Jun Sakuma

Feature allocation models generalize species sampling models by allowing every observation to belong to more than one species, now called features. Under the popular Bernoulli product model for feature allocation, given $n$ samples, we…

Statistics Theory · Mathematics 2020-09-22 Fadhel Ayed , Marco Battiston , Federico Camerlenghi , Stefano Favaro

A random variable is sampled from a discrete distribution. The missing mass is the probability of the set of points not observed in the sample. We sharpen and simplify McAllester and Ortiz's results (JMLR, 2003) bounding the probability of…

Probability · Mathematics 2012-10-12 Daniel Berend , Aryeh Kontorovich

We consider the problem of estimating the $L_1$ distance between two discrete probability measures $P$ and $Q$ from empirical data in a nonasymptotic and large alphabet setting. When $Q$ is known and one obtains $n$ samples from $P$, we…

Statistics Theory · Mathematics 2018-06-26 Jiantao Jiao , Yanjun Han , Tsachy Weissman

The problem of estimating the missing mass or total probability of unseen elements in a sequence of $n$ random samples is considered under the squared error loss function. The worst-case risk of the popular Good-Turing estimator is shown to…

Information Theory · Computer Science 2017-05-16 Nikhilesh Rajaraman , Andrew Thangaraj , Ananda Theertha Suresh

By a mixture density is meant a density of the form $\pi_{\mu}(\cdot)=\int\pi_{\theta}(\cdot)\times\mu(d\theta)$, where $(\pi_{\theta})_{\theta\in\Theta}$ is a family of probability densities and $\mu$ is a probability measure on $\Theta$.…

Statistics Theory · Mathematics 2016-08-16 François Roueff , Tobias Rydén

We address a parametric joint detection-estimation problem for discrete signals of the form $x(t) = \sum_{n=1}^{N} \alpha_n e^{-i \lambda_n t } + \epsilon_t$, $t \in \mathbb{N}$, with an additive noise represented by independent centered…

Classical Analysis and ODEs · Mathematics 2018-08-14 Illya M. Karabash , Jürgen Prestin

The empirical distribution function assigns mass $1/n$ to each of the $n$ observations in a sample. As these are highly variable, estimation error may be reduced by replacing them with estimated observations that are asymptotically less…

Methodology · Statistics 2026-05-26 Tommaso Lando , Lorenzo Tedesco

The Good-Turing (GT) estimator for the missing mass (i.e., total probability of missing symbols) in $n$ samples is the number of symbols that appeared exactly once divided by $n$. For i.i.d. samples, the bias and squared-error risk of the…

Information Theory · Computer Science 2023-05-30 Prafulla Chandra , Andrew Thangaraj , Nived Rajaraman

We study the problem of estimating a distribution over a finite alphabet from an i.i.d. sample, with accuracy measured in relative entropy (Kullback-Leibler divergence). While optimal bounds on the expected risk are known, high-probability…

Statistics Theory · Mathematics 2026-02-27 Jaouad Mourtada

The paper suggests a simple method of deriving minimax lower bounds to the accuracy of statistical inference on heavy tails. A well-known result by Hall and Welsh (Ann. Statist. 12 (1984) 1079-1084) states that if $\hat{\alpha}_n$ is an…

Statistics Theory · Mathematics 2014-03-14 S. Y. Novak
‹ Prev 1 2 3 10 Next ›