English

On Focal Loss for Class-Posterior Probability Estimation: A Theoretical Perspective

Machine Learning 2020-12-15 v2 Machine Learning

Abstract

The focal loss has demonstrated its effectiveness in many real-world applications such as object detection and image classification, but its theoretical understanding has been limited so far. In this paper, we first prove that the focal loss is classification-calibrated, i.e., its minimizer surely yields the Bayes-optimal classifier and thus the use of the focal loss in classification can be theoretically justified. However, we also prove a negative fact that the focal loss is not strictly proper, i.e., the confidence score of the classifier obtained by focal loss minimization does not match the true class-posterior probability and thus it is not reliable as a class-posterior probability estimator. To mitigate this problem, we next prove that a particular closed-form transformation of the confidence score allows us to recover the true class-posterior probability. Through experiments on benchmark datasets, we demonstrate that our proposed transformation significantly improves the accuracy of class-posterior probability estimation.

Keywords

Cite

@article{arxiv.2011.09172,
  title  = {On Focal Loss for Class-Posterior Probability Estimation: A Theoretical Perspective},
  author = {Nontawat Charoenphakdee and Jayakorn Vongkulbhisal and Nuttapong Chairatanakul and Masashi Sugiyama},
  journal= {arXiv preprint arXiv:2011.09172},
  year   = {2020}
}

Comments

57 pages

R2 v1 2026-06-23T20:20:26.606Z