English

DeepFaceFlow: In-the-wild Dense 3D Facial Motion Estimation

Computer Vision and Pattern Recognition 2020-05-18 v1

Abstract

Dense 3D facial motion capture from only monocular in-the-wild pairs of RGB images is a highly challenging problem with numerous applications, ranging from facial expression recognition to facial reenactment. In this work, we propose DeepFaceFlow, a robust, fast, and highly-accurate framework for the dense estimation of 3D non-rigid facial flow between pairs of monocular images. Our DeepFaceFlow framework was trained and tested on two very large-scale facial video datasets, one of them of our own collection and annotation, with the aid of occlusion-aware and 3D-based loss function. We conduct comprehensive experiments probing different aspects of our approach and demonstrating its improved performance against state-of-the-art flow and 3D reconstruction methods. Furthermore, we incorporate our framework in a full-head state-of-the-art facial video synthesis method and demonstrate the ability of our method in better representing and capturing the facial dynamics, resulting in a highly-realistic facial video synthesis. Given registered pairs of images, our framework generates 3D flow maps at ~60 fps.

Keywords

Cite

@article{arxiv.2005.07298,
  title  = {DeepFaceFlow: In-the-wild Dense 3D Facial Motion Estimation},
  author = {Mohammad Rami Koujan and Anastasios Roussos and Stefanos Zafeiriou},
  journal= {arXiv preprint arXiv:2005.07298},
  year   = {2020}
}

Comments

to be published in the IEEE conference on Computer Vision and Pattern Recognition (CVPR). 2020

R2 v1 2026-06-23T15:33:44.625Z