English

Distributionally Robust Losses for Latent Covariate Mixtures

Machine Learning 2022-08-12 v2 Machine Learning

Abstract

While modern large-scale datasets often consist of heterogeneous subpopulations -- for example, multiple demographic groups or multiple text corpora -- the standard practice of minimizing average loss fails to guarantee uniformly low losses across all subpopulations. We propose a convex procedure that controls the worst-case performance over all subpopulations of a given size. Our procedure comes with finite-sample (nonparametric) convergence guarantees on the worst-off subpopulation. Empirically, we observe on lexical similarity, wine quality, and recidivism prediction tasks that our worst-case procedure learns models that do well against unseen subpopulations.

Keywords

Cite

@article{arxiv.2007.13982,
  title  = {Distributionally Robust Losses for Latent Covariate Mixtures},
  author = {John Duchi and Tatsunori Hashimoto and Hongseok Namkoong},
  journal= {arXiv preprint arXiv:2007.13982},
  year   = {2022}
}

Comments

First released in 2019 on a personal website; published in Operations Research in 2022