English
Related papers

Related papers: Towards Expressive Video Dubbing with Multiscale M…

200 papers

We introduce a video framework for modeling the association between verbal and non-verbal communication during dyadic conversation. Given the input speech of a speaker, our approach retrieves a video of a listener, who has facial…

Computer Vision and Pattern Recognition · Computer Science 2023-01-27 Scott Geng , Revant Teotia , Purva Tendulkar , Sachit Menon , Carl Vondrick

Movie Dubbing aims to convert scripts into speeches that align with the given movie clip in both temporal and emotional aspects while preserving the vocal timbre of a given brief reference audio. Existing methods focus primarily on reducing…

Video captioning is a challenging task that captures different visual parts and describes them in sentences, for it requires visual and linguistic coherence. The attention mechanism in the current video captioning method learns to assign…

Computer Vision and Pattern Recognition · Computer Science 2021-10-19 Zhixin Sun , Xian Zhong , Shuqin Chen , Lin Li , Luo Zhong

Visual dubbing, the synchronization of facial movements with new speech, is crucial for making content accessible across different languages, enabling broader global reach. However, current methods face significant limitations. Existing…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Binyamin Manela , Sharon Gannot , Ethan Fetyaya

Multimodal dialogue emotion recognition captures emotional cues by fusing text, visual, and audio modalities. However, existing approaches still suffer from notable limitations in modeling emotional dependencies and learning multimodal…

Multimedia · Computer Science 2026-03-12 Yunsheng Wang , Yuntao Shou , Yilong Tan , Wei Ai , Tao Meng , Keqin Li

While most conversational AI systems focus on textual dialogue only, conditioning utterances on visual context (when it's available) can lead to more realistic conversations. Unfortunately, a major challenge for incorporating visual context…

Computer Vision and Pattern Recognition · Computer Science 2021-03-30 Paul Hongsuck Seo , Arsha Nagrani , Cordelia Schmid

Speech-driven 3D facial animation has improved a lot recently while most related works only utilize acoustic modality and neglect the influence of visual and textual cues, leading to unsatisfactory results in terms of precision and…

Computer Vision and Pattern Recognition · Computer Science 2023-12-06 Tianshun Han , Shengnan Gui , Yiqing Huang , Baihui Li , Lijian Liu , Benjia Zhou , Ning Jiang , Quan Lu , Ruicong Zhi , Yanyan Liang , Du Zhang , Jun Wan

In this work, we focus on leveraging facial cues beyond the lip region for robust Audio-Visual Speech Enhancement (AVSE). The facial region, encompassing the lip region, reflects additional speech-related attributes such as gender, skin…

Computer Vision and Pattern Recognition · Computer Science 2023-11-27 Feixiang Wang , Shuang Yang , Shiguang Shan , Xilin Chen

We present Empathic Prompting, a novel framework for multimodal human-AI interaction that enriches Large Language Model (LLM) conversations with implicit non-verbal context. The system integrates a commercial facial expression recognition…

Human-Computer Interaction · Computer Science 2026-05-29 Lorenzo Stacchio , Andrea Ubaldi , Alessandro Galdelli , Maurizio Mauri , Emanuele Frontoni , Andrea Gaggioli

We present VoiceCraft-Dub, a novel approach for automated video dubbing that synthesizes high-quality speech from text and facial cues. This task has broad applications in filmmaking, multimedia creation, and assisting voice-impaired…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Kim Sung-Bin , Jeongsoo Choi , Puyuan Peng , Joon Son Chung , Tae-Hyun Oh , David Harwath

Understanding social interaction in video requires reasoning over a dynamic interplay of verbal and non-verbal cues: who is speaking, to whom, and with what gaze or gestures. While Multimodal Large Language Models (MLLMs) are natural…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Liangyang Ouyang , Yifei Huang , Mingfang Zhang , Caixin Kang , Ryosuke Furuta , Yoichi Sato

This paper focuses on designing a noise-robust end-to-end Audio-Visual Speech Recognition (AVSR) system. To this end, we propose Visual Context-driven Audio Feature Enhancement module (V-CAFE) to enhance the input noisy audio speech with a…

Sound · Computer Science 2022-07-14 Joanna Hong , Minsu Kim , Daehun Yoo , Yong Man Ro

Automatic video commentary systems are widely used on multimedia social media platforms to extract factual information about video content. However, current systems may overlook essential para-linguistic cues, including emotion and…

Human-Computer Interaction · Computer Science 2025-06-23 Qixin Wang , Songtao Zhou , Zeyu Jin , Chenglin Guo , Shikun Sun , Xiaoyu Qin

Human affective behavior analysis has received much attention in human-computer interaction (HCI). In this paper, we introduce our submission to the CVPR 2022 Competition on Affective Behavior Analysis in-the-wild (ABAW). To fully exploit…

Computer Vision and Pattern Recognition · Computer Science 2022-04-19 Wei Zhang , Feng Qiu , Suzhen Wang , Hao Zeng , Zhimeng Zhang , Rudong An , Bowen Ma , Yu Ding

Automatic speech-based affect recognition of individuals in dyadic conversation is a challenging task, in part because of its heavy reliance on manual pre-processing. Traditional approaches frequently require hand-crafted speech features…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-24 Huili Chen , Yue Zhang , Felix Weninger , Rosalind Picard , Cynthia Breazeal , Hae Won Park

Multimodal sentiment analysis (MSA), which supposes to improve text-based sentiment analysis with associated acoustic and visual modalities, is an emerging research area due to its potential applications in Human-Computer Interaction (HCI).…

Multimedia · Computer Science 2022-09-07 Yihe Liu , Ziqi Yuan , Huisheng Mao , Zhiyun Liang , Wanqiuyue Yang , Yuanzhe Qiu , Tie Cheng , Xiaoteng Li , Hua Xu , Kai Gao

Multimodal analysis has recently drawn much interest in affective computing, since it can improve the overall accuracy of emotion recognition over isolated uni-modal approaches. The most effective techniques for multimodal emotion…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 R. Gnana Praveen , Eric Granger , Patrick Cardinal

In this paper we present VDTTS, a Visually-Driven Text-to-Speech model. Motivated by dubbing, VDTTS takes advantage of video frames as an additional input alongside text, and generates speech that matches the video signal. We demonstrate…

Computer Vision and Pattern Recognition · Computer Science 2022-03-25 Michael Hassid , Michelle Tadmor Ramanovich , Brendan Shillingford , Miaosen Wang , Ye Jia , Tal Remez

We propose Context-Adaptive Multi-Prompt Embedding, a novel approach to enrich semantic representations in vision-language contrastive learning. Unlike standard CLIP-style models that rely on a single text embedding, our method introduces…

Machine Learning · Computer Science 2025-08-07 Dahun Kim , Anelia Angelova