English

A Realistic Face-to-Face Conversation System based on Deep Neural Networks

Computer Vision and Pattern Recognition 2019-08-22 v1 Sound Audio and Speech Processing Image and Video Processing

Abstract

To improve the experiences of face-to-face conversation with avatar, this paper presents a novel conversation system. It is composed of two sequence-to-sequence models respectively for listening and speaking and a Generative Adversarial Network (GAN) based realistic avatar synthesizer. The models exploit the facial action and head pose to learn natural human reactions. Based on the models' output, the synthesizer uses the Pixel2Pixel model to generate realistic facial images. To show the improvement of our system, we use a 3D model based avatar driving scheme as a reference. We train and evaluate our neural networks with the data from ESPN shows. Experimental results show that our conversation system can generate natural facial reactions and realistic facial images.

Keywords

Cite

@article{arxiv.1908.07750,
  title  = {A Realistic Face-to-Face Conversation System based on Deep Neural Networks},
  author = {Zezhou Chen and Zhaoxiang Liu and Huan Hu and Jinqiang Bai and Shiguo Lian and Fuyuan Shi and Kai Wang},
  journal= {arXiv preprint arXiv:1908.07750},
  year   = {2019}
}

Comments

Accepted to ICCV 2019 workshop

R2 v1 2026-06-23T10:52:58.677Z