English

The Rate of Convergence of AdaBoost

Optimization and Control 2011-06-30 v1 Artificial Intelligence Machine Learning

Abstract

The AdaBoost algorithm was designed to combine many "weak" hypotheses that perform slightly better than random guessing into a "strong" hypothesis that has very low error. We study the rate at which AdaBoost iteratively converges to the minimum of the "exponential loss." Unlike previous work, our proofs do not require a weak-learning assumption, nor do they require that minimizers of the exponential loss are finite. Our first result shows that at iteration tt, the exponential loss of AdaBoost's computed parameter vector will be at most ϵ\epsilon more than that of any parameter vector of 1\ell_1-norm bounded by BB in a number of rounds that is at most a polynomial in BB and 1/ϵ1/\epsilon. We also provide lower bounds showing that a polynomial dependence on these parameters is necessary. Our second result is that within C/ϵC/\epsilon iterations, AdaBoost achieves a value of the exponential loss that is at most ϵ\epsilon more than the best possible value, where CC depends on the dataset. We show that this dependence of the rate on ϵ\epsilon is optimal up to constant factors, i.e., at least Ω(1/ϵ)\Omega(1/\epsilon) rounds are necessary to achieve within ϵ\epsilon of the optimal exponential loss.

Keywords

Cite

@article{arxiv.1106.6024,
  title  = {The Rate of Convergence of AdaBoost},
  author = {Indraneel Mukherjee and Cynthia Rudin and Robert E. Schapire},
  journal= {arXiv preprint arXiv:1106.6024},
  year   = {2011}
}

Comments

A preliminary version will appear in COLT 2011