English

Exact Exponent in Optimal Rates for Crowdsourcing

Machine Learning 2016-05-27 v2 Statistics Theory Statistics Theory

Abstract

In many machine learning applications, crowdsourcing has become the primary means for label collection. In this paper, we study the optimal error rate for aggregating labels provided by a set of non-expert workers. Under the classic Dawid-Skene model, we establish matching upper and lower bounds with an exact exponent mI(π)mI(\pi) in which mm is the number of workers and I(π)I(\pi) the average Chernoff information that characterizes the workers' collective ability. Such an exact characterization of the error exponent allows us to state a precise sample size requirement m>1I(π)log1ϵm>\frac{1}{I(\pi)}\log\frac{1}{\epsilon} in order to achieve an ϵ\epsilon misclassification error. In addition, our results imply the optimality of various EM algorithms for crowdsourcing initialized by consistent estimators.

Keywords

Cite

@article{arxiv.1605.07696,
  title  = {Exact Exponent in Optimal Rates for Crowdsourcing},
  author = {Chao Gao and Yu Lu and Dengyong Zhou},
  journal= {arXiv preprint arXiv:1605.07696},
  year   = {2016}
}

Comments

To appear in the Proceedings of the 33rd International Conference on Machine Learning, New York, NY, USA, 2016

R2 v1 2026-06-22T14:08:51.680Z