English

Adaptive Estimation of Shannon Entropy

Information Theory 2019-01-03 v2 math.IT

Abstract

We consider estimating the Shannon entropy of a discrete distribution PP from nn i.i.d. samples. Recently, Jiao, Venkat, Han, and Weissman, and Wu and Yang constructed approximation theoretic estimators that achieve the minimax L2L_2 rates in estimating entropy. Their estimators are consistent given nSlnSn \gg \frac{S}{\ln S} samples, where SS is the alphabet size, and it is the best possible sample complexity. In contrast, the Maximum Likelihood Estimator (MLE), which is the empirical entropy, requires nSn\gg S samples. In the present paper we significantly refine the minimax results of existing work. To alleviate the pessimism of minimaxity, we adopt the adaptive estimation framework, and show that the minimax rate-optimal estimator in Jiao, Venkat, Han, and Weissman achieves the minimax rates simultaneously over a nested sequence of subsets of distributions PP, without knowing the alphabet size SS or which subset PP lies in. In other words, their estimator is adaptive with respect to this nested sequence of the parameter space, which is characterized by the entropy of the distribution. We also characterize the maximum risk of the MLE over this nested sequence, and show, for every subset in the sequence, that the performance of the minimax rate-optimal estimator with nn samples is essentially that of the MLE with nlnnn\ln n samples, thereby further substantiating the generality of the phenomenon identified by Jiao, Venkat, Han, and Weissman.

Keywords

Cite

@article{arxiv.1502.00326,
  title  = {Adaptive Estimation of Shannon Entropy},
  author = {Yanjun Han and Jiantao Jiao and Tsachy Weissman},
  journal= {arXiv preprint arXiv:1502.00326},
  year   = {2019}
}
R2 v1 2026-06-22T08:18:24.790Z