English

Unsupervised Deep Representations for Learning Audience Facial Behaviors

Computer Vision and Pattern Recognition 2018-05-14 v1

Abstract

In this paper, we present an unsupervised learning approach for analyzing facial behavior based on a deep generative model combined with a convolutional neural network (CNN). We jointly train a variational auto-encoder (VAE) and a generative adversarial network (GAN) to learn a powerful latent representation from footage of audiences viewing feature-length movies. We show that the learned latent representation successfully encodes meaningful signatures of behaviors related to audience engagement (smiling & laughing) and disengagement (yawning). Our results provide a proof of concept for a more general methodology for annotating hard-to-label multimedia data featuring sparse examples of signals of interest.

Keywords

Cite

@article{arxiv.1805.04136,
  title  = {Unsupervised Deep Representations for Learning Audience Facial Behaviors},
  author = {Suman Saha and Rajitha Navarathna and Leonhard Helminger and Romann Weber},
  journal= {arXiv preprint arXiv:1805.04136},
  year   = {2018}
}
R2 v1 2026-06-23T01:51:24.832Z