中文
相关论文

相关论文: EAD-Net: Emotion-Aware Talking Head Generation wit…

200 篇论文

Sarcasm Explanation in Dialogue (SED) is a new yet challenging task, which aims to generate a natural language explanation for the given sarcastic dialogue that involves multiple modalities (\ie utterance, video, and audio). Although…

计算与语言 · 计算机科学 2025-01-07 Kun Ouyang , Liqiang Jing , Xuemeng Song , Meng Liu , Yupeng Hu , Liqiang Nie

Recently, physiological data such as electroencephalography (EEG) signals have attracted significant attention in affective computing. In this context, the main goal is to design an automated model that can assess emotional states. Lately,…

机器学习 · 计算机科学 2023-07-07 Shadi Sartipi , Mastaneh Torkamani-Azar , Mujdat Cetin

Recent advancements in video diffusion models have significantly enhanced audio-driven portrait animation. However, current methods still suffer from flickering, identity drift, and poor audio-visual synchronization. These issues primarily…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Zhenjie Liu , Jianzhang Lu , Renjie Lu , Cong Liang , Shangfei Wang

Given a piece of text, a video clip, and a reference audio, the movie dubbing task aims to generate speech that aligns with the video while cloning the desired voice. The existing methods have two primary deficiencies: (1) They struggle to…

Real-world talking faces often accompany with natural head movement. However, most existing talking face video generation methods only consider facial animation with fixed head pose. In this paper, we address this problem by proposing a…

计算机视觉与模式识别 · 计算机科学 2020-03-06 Ran Yi , Zipeng Ye , Juyong Zhang , Hujun Bao , Yong-Jin Liu

Creating a realistic animatable avatar from a single static portrait remains challenging. Existing approaches often struggle to capture subtle facial expressions, the associated global body movements, and the dynamic background. To address…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Mengchao Wang , Qiang Wang , Fan Jiang , Yaqi Fan , Yunpeng Zhang , Yonggang Qi , Kun Zhao , Mu Xu

Speech-driven 3D face animation technique, extending its applications to various multimedia fields. Previous research has generated promising realistic lip movements and facial expressions from audio signals. However, traditional regression…

计算机视觉与模式识别 · 计算机科学 2023-08-31 Ziqiao Peng , Yihao Luo , Yue Shi , Hao Xu , Xiangyu Zhu , Jun He , Hongyan Liu , Zhaoxin Fan

Generating realistic talking-head videos remains challenging due to persistent issues such as imperfect lip synchronization, unnatural motion, and evaluation metrics that correlate poorly with human perception. We propose FlowPortrait, a…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Weiting Tan , Andy T. Liu , Ming Tu , Xinghua Qu , Philipp Koehn , Lu Lu

Empathetic Response Generation (ERG) is one of the key tasks of the affective computing area, which aims to produce emotionally nuanced and compassionate responses to user's queries. However, existing ERG research is predominantly confined…

多媒体 · 计算机科学 2025-02-10 Han Zhang , Zixiang Meng , Meng Luo , Hong Han , Lizi Liao , Erik Cambria , Hao Fei

Generating talking avatar driven by audio remains a significant challenge. Existing methods typically require high computational costs and often lack sufficient facial detail and realism, making them unsuitable for applications that demand…

计算机视觉与模式识别 · 计算机科学 2025-01-27 Yujian Liu , Shidang Xu , Jing Guo , Dingbin Wang , Zairan Wang , Xianfeng Tan , Xiaoli Liu

Multimodal emotion recognition (MER) is crucial for enabling emotionally intelligent systems that perceive and respond to human emotions. However, existing methods suffer from limited cross-modal interaction and imbalanced contributions…

多媒体 · 计算机科学 2025-07-30 Zeyu Deng , Yanhui Lu , Jiashu Liao , Shuang Wu , Chongfeng Wei

In the task of talking face generation, the objective is to generate a face video with lips synchronized to the corresponding audio while preserving visual details and identity information. Current methods face the challenge of learning…

计算机视觉与模式识别 · 计算机科学 2024-05-08 Dogucan Yaman , Fevziye Irem Eyiokur , Leonard Bärmann , Seymanur Aktı , Hazım Kemal Ekenel , Alexander Waibel

Understanding learner emotions in online education is critical for improving engagement and personalized instruction. While prior work in emotion recognition has explored multimodal fusion and temporal modeling, existing methods often rely…

机器学习 · 计算机科学 2025-10-13 S M Rafiuddin

Recent diffusion-based talking face generation models have demonstrated impressive potential in synthesizing videos that accurately match a speech audio clip with a given reference identity. However, existing approaches still encounter…

计算机视觉与模式识别 · 计算机科学 2025-10-16 Xingpei Ma , Jiaran Cai , Yuansheng Guan , Shenneng Huang , Qiang Zhang , Shunsi Zhang

The creation of increasingly vivid 3D talking face has become a hot topic in recent years. Currently, most speech-driven works focus on lip synchronisation but neglect to effectively capture the correlations between emotions and facial…

计算机视觉与模式识别 · 计算机科学 2025-05-19 Yihong Lin , Liang Peng , Zhaoxin Fan , Xianjia Wu , Jianqiao Hu , Xiandong Li , Wenxiong Kang , Songju Lei

Diffusion models have shown impressive potential on talking head generation. While plausible appearance and talking effect are achieved, these methods still suffer from temporal, 3D or expression inconsistency due to the error accumulation…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Haijie Yang , Zhenyu Zhang , Hao Tang , Jianjun Qian , Jian Yang

While existing one-shot talking head generation models have achieved progress in coarse-grained emotion editing, there is still a lack of fine-grained emotion editing models with high interpretability. We argue that for an approach to be…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Guanwen Feng , Zhihao Qian , Yunan Li , Siyu Jin , Qiguang Miao , Chi-Man Pun

Over the last few decades, many aspects of human life have been enhanced with virtual domains, from the advent of digital assistants such as Amazon's Alexa and Apple's Siri to the latest metaverse efforts of the rebranded Meta. These trends…

计算机视觉与模式识别 · 计算机科学 2023-03-27 Siddarth Ravichandran , Ondřej Texler , Dimitar Dinev , Hyun Jae Kang

We present a multimodal learning-based method to simultaneously synthesize co-speech facial expressions and upper-body gestures for digital characters using RGB video data captured using commodity cameras. Our approach learns from sparse…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Uttaran Bhattacharya , Aniket Bera , Dinesh Manocha

Generating vivid and diverse 3D co-speech gestures is crucial for various applications in animating virtual avatars. While most existing methods can generate gestures from audio directly, they usually overlook that emotion is one of the key…

计算机视觉与模式识别 · 计算机科学 2024-01-04 Xingqun Qi , Chen Liu , Lincheng Li , Jie Hou , Haoran Xin , Xin Yu