English

Suboptimality of Penalized Empirical Risk Minimization in Classification

Statistics Theory 2008-12-02 v1 Risk Management Statistics Theory

Abstract

Let \cF\cF be a set of MM classification procedures with values in [1,1][-1,1]. Given a loss function, we want to construct a procedure which mimics at the best possible rate the best procedure in \cF\cF. This fastest rate is called optimal rate of aggregation. Considering a continuous scale of loss functions with various types of convexity, we prove that optimal rates of aggregation can be either ((logM)/n)1/2((\log M)/n)^{1/2} or (logM)/n(\log M)/n. We prove that, if all the MM classifiers are binary, the (penalized) Empirical Risk Minimization procedures are suboptimal (even under the margin/low noise condition) when the loss function is somewhat more than convex, whereas, in that case, aggregation procedures with exponential weights achieve the optimal rate of aggregation.

Keywords

Cite

@article{arxiv.math/0703811,
  title  = {Suboptimality of Penalized Empirical Risk Minimization in Classification},
  author = {Guillaume Lecué},
  journal= {arXiv preprint arXiv:math/0703811},
  year   = {2008}
}

Comments

15 pages

R2 v1 2026-07-22T17:53:18.578Z