English

HUMBO: Bridging Response Generation and Facial Expression Synthesis

Computation and Language 2021-09-01 v3 Computer Vision and Pattern Recognition Multimedia

Abstract

Spoken dialogue systems that assist users to solve complex tasks such as movie ticket booking have become an emerging research topic in artificial intelligence and natural language processing areas. With a well-designed dialogue system as an intelligent personal assistant, people can accomplish certain tasks more easily via natural language interactions. Today there are several virtual intelligent assistants in the market; however, most systems only focus on textual or vocal interaction. In this paper, we present HUMBO, a system aiming at generating dialogue responses and simultaneously synthesize corresponding visual expressions on faces for better multimodal interaction. HUMBO can (1) let users determine the appearances of virtual assistants by a single image, and (2) generate coherent emotional utterances and facial expressions on the user-provided image. This is not only a brand new research direction but more importantly, an ultimate step toward more human-like virtual assistants.

Keywords

Cite

@article{arxiv.1905.11240,
  title  = {HUMBO: Bridging Response Generation and Facial Expression Synthesis},
  author = {Shang-Yu Su and Po-Wei Lin and Yun-Nung Chen},
  journal= {arXiv preprint arXiv:1905.11240},
  year   = {2021}
}

Comments

The first two authors contributed to this work equally

R2 v1 2026-06-23T09:26:40.209Z