English
Related papers

Related papers: Affective Faces for Goal-Driven Dyadic Communicati…

200 papers

Effective human behavior modeling is critical for successful human-robot interaction. Current state-of-the-art approaches for predicting listening head behavior during dyadic conversations employ continuous-to-discrete representations,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-09 Tri Tung Nguyen Nguyen , Quang Tien Dam , Dinh Tuan Tran , Joo-Ho Lee

This article presents a multimodal emotion recognition module integrated into a proactive Socially Interactive Agent (SIA) powered by generative artificial intelligence. The system evaluates real-time affective states through two distinct…

Human-Computer Interaction · Computer Science 2026-05-21 Adnana Dragut , Raquel Lacuesta , F. Xavier Gaya-Morey , Jose M. Buades-Rubio

In this paper, we present a methodology for the development of embodied conversational agents for social virtual worlds. The agents provide multimodal communication with their users in which speech interaction is included. Our proposal…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-29 D. Griol , A. Sanchis , J. M. Molina , Z. Callejas

In order to build self-consistent personalized dialogue agents, previous research has mostly focused on textual persona that delivers personal facts or personalities. However, to fully describe the multi-faceted nature of persona, image…

Computation and Language · Computer Science 2023-05-30 Jaewoo Ahn , Yeda Song , Sangdoo Yun , Gunhee Kim

The development of virtual agents has enabled human-avatar interactions to become increasingly rich and varied. Moreover, an expressive virtual agent i.e. that mimics the natural expression of emotions, enhances social interaction between a…

Computer Vision and Pattern Recognition · Computer Science 2022-05-23 Hugo Bohy , Ahmad Hammoudeh , Antoine Maiorca , Stéphane Dupont , Thierry Dutoit

To facilitate the research on intelligent and human-like chatbots with multi-modal context, we introduce a new video-based multi-modal dialogue dataset, called TikTalk. We collect 38K videos from a popular video-sharing platform, along with…

Computation and Language · Computer Science 2023-09-11 Hongpeng Lin , Ludan Ruan , Wenke Xia , Peiyu Liu , Jingyuan Wen , Yixin Xu , Di Hu , Ruihua Song , Wayne Xin Zhao , Qin Jin , Zhiwu Lu

While considerable progress has been made in achieving accurate lip synchronization for 3D speech-driven talking face generation, the task of incorporating expressive facial detail synthesis aligned with the speaker's speaking status…

Computer Vision and Pattern Recognition · Computer Science 2024-02-27 Yasheng Sun , Wenqing Chu , Hang Zhou , Kaisiyuan Wang , Hideki Koike

Goal-oriented conversational agents are becoming prevalent in our daily lives. For these systems to engage users and achieve their goals, they need to exhibit appropriate social behavior as well as provide informative replies that guide…

Computation and Language · Computer Science 2021-01-01 Yi-Chia Wang , Alexandros Papangelis , Runze Wang , Zhaleh Feizollahi , Gokhan Tur , Robert Kraut

Chatbots via large language models (LLMs) generate fluent responses but often struggle with when to speak, especially for brief, timely listener reactions during ongoing dialogue. We present a multimodal strategy for LLMs, which leverages…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Zikai Liao , Yi Ouyang , Yi-Lun Lee , Chen-Ping Yu , Yi-Hsuan Tsai , Zhaozheng Yin

In this paper, we propose a talking face generation method that takes an audio signal as input and a short target video clip as reference, and synthesizes a photo-realistic video of the target face with natural lip motions, head poses, and…

Computer Vision and Pattern Recognition · Computer Science 2021-08-19 Chenxu Zhang , Yifan Zhao , Yifei Huang , Ming Zeng , Saifeng Ni , Madhukar Budagavi , Xiaohu Guo

Given an arbitrary face image and an arbitrary speech clip, the proposed work attempts to generating the talking face video with accurate lip synchronization while maintaining smooth transition of both lip and facial movement over the…

Computer Vision and Pattern Recognition · Computer Science 2019-07-29 Yang Song , Jingwen Zhu , Dawei Li , Xiaolong Wang , Hairong Qi

We introduce FaceTalk, a novel generative approach designed for synthesizing high-fidelity 3D motion sequences of talking human heads from input audio signal. To capture the expressive, detailed nature of human heads, including hair, ears,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Shivangi Aneja , Justus Thies , Angela Dai , Matthias Nießner

The paper introduces AniTalker, an innovative framework designed to generate lifelike talking faces from a single portrait. Unlike existing models that primarily focus on verbal cues such as lip synchronization and fail to capture the…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Tao Liu , Feilong Chen , Shuai Fan , Chenpeng Du , Qi Chen , Xie Chen , Kai Yu

Understanding social interaction in video requires reasoning over a dynamic interplay of verbal and non-verbal cues: who is speaking, to whom, and with what gaze or gestures. While Multimodal Large Language Models (MLLMs) are natural…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Liangyang Ouyang , Yifei Huang , Mingfang Zhang , Caixin Kang , Ryosuke Furuta , Yoichi Sato

Audio-driven 3D facial animation has been widely explored, but achieving realistic, human-like performance is still unsolved. This is due to the lack of available 3D datasets, models, and standard evaluation metrics. To address this, we…

Computer Vision and Pattern Recognition · Computer Science 2019-05-09 Daniel Cudeiro , Timo Bolkart , Cassidy Laidlaw , Anurag Ranjan , Michael J. Black

Human interactions are characterized by explicit as well as implicit channels of communication. While the explicit channel transmits overt messages, the implicit ones transmit hidden messages about the communicator (e.g., his/her intentions…

Human-Computer Interaction · Computer Science 2017-07-12 Katie Hoemann , Behnaz Rezaei , Stacy C. Marsella , Sarah Ostadabbas

For audio-driven visual dubbing, it remains a considerable challenge to uphold and highlight speaker's persona while synthesizing accurate lip synchronization. Existing methods fall short of capturing speaker's unique speaking style or…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Longhao Zhang , Shuang Liang , Zhipeng Ge , Tianshu Hu

It is in high demand to generate facial animation with high realism, but it remains a challenging task. Existing approaches of speech-driven facial animation can produce satisfactory mouth movement and lip synchronization, but show weakness…

Computer Vision and Pattern Recognition · Computer Science 2025-05-08 Yutong Chen , Junhong Zhao , Wei-Qiang Zhang

Speech emotion recognition is a challenging problem because human convey emotions in subtle and complex ways. For emotion recognition on human speech, one can either extract emotion related features from audio signals or employ speech…

Computation and Language · Computer Science 2020-04-06 Haiyang Xu , Hui Zhang , Kun Han , Yun Wang , Yiping Peng , Xiangang Li

Gaze following estimates gaze targets of in-scene person by understanding human behavior and scene information. Existing methods usually analyze scene images for gaze following. However, compared with visual images, audio also provides…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Yuqi Hou , Zhongqun Zhang , Nora Horanyi , Jaewon Moon , Yihua Cheng , Hyung Jin Chang