English

On the use of automatically generated synthetic image datasets for benchmarking face recognition

Computer Vision and Pattern Recognition 2021-10-04 v1

Abstract

The availability of large-scale face datasets has been key in the progress of face recognition. However, due to licensing issues or copyright infringement, some datasets are not available anymore (e.g. MS-Celeb-1M). Recent advances in Generative Adversarial Networks (GANs), to synthesize realistic face images, provide a pathway to replace real datasets by synthetic datasets, both to train and benchmark face recognition (FR) systems. The work presented in this paper provides a study on benchmarking FR systems using a synthetic dataset. First, we introduce the proposed methodology to generate a synthetic dataset, without the need for human intervention, by exploiting the latent structure of a StyleGAN2 model with multiple controlled factors of variation. Then, we confirm that (i) the generated synthetic identities are not data subjects from the GAN's training dataset, which is verified on a synthetic dataset with 10K+ identities; (ii) benchmarking results on the synthetic dataset are a good substitution, often providing error rates and system ranking similar to the benchmarking on the real dataset.

Keywords

Cite

@article{arxiv.2106.04215,
  title  = {On the use of automatically generated synthetic image datasets for benchmarking face recognition},
  author = {Laurent Colbois and Tiago de Freitas Pereira and Sébastien Marcel},
  journal= {arXiv preprint arXiv:2106.04215},
  year   = {2021}
}

Comments

11 pages, Accepted for publication in the 2021 International Joint Conference on Biometrics (IJCB 2021)