中文
相关论文

相关论文: Disentangle Identity, Cooperate Emotion: Correlati…

200 篇论文

Portrait animation from a single source image and a driving video is a long-standing problem. Recent approaches tend to adopt diffusion-based image/video generation models for realistic and expressive animation. However, none of these…

计算机视觉与模式识别 · 计算机科学 2025-12-18 Yuxiang Shi , Zhe Li , Yanwen Wang , Hao Zhu , Xun Cao , Ligang Liu

The domain of 3D talking head generation has witnessed significant progress in recent years. A notable challenge in this field consists in blending speech-related motions with expression dynamics, which is primarily caused by the lack of…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Federico Nocentini , Claudio Ferrari , Stefano Berretti

Speech-driven 3D facial animation seeks to produce lifelike facial expressions that are synchronized with the speech content and its emotional nuances, finding applications in various multimedia fields. However, previous methods often…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Yixuan Zhang , Qing Chang , Yuxi Wang , Guang Chen , Zhaoxiang Zhang , Junran Peng

Audio-driven single-image talking portrait generation plays a crucial role in virtual reality, digital human creation, and filmmaking. Existing approaches are generally categorized into keypoint-based and image-based methods. Keypoint-based…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Chaolong Yang , Kai Yao , Yuyao Yan , Chenru Jiang , Weiguang Zhao , Jie Sun , Guangliang Cheng , Yifei Zhang , Bin Dong , Kaizhu Huang

Multimodal-driven talking face generation refers to animating a portrait with the given pose, expression, and gaze transferred from the driving image and video, or estimated from the text and audio. However, existing methods ignore the…

计算机视觉与模式识别 · 计算机科学 2023-05-10 Chao Xu , Shaoting Zhu , Junwei Zhu , Tianxin Huang , Jiangning Zhang , Ying Tai , Yong Liu

The successful emotional conversation system depends on sufficient perception and appropriate expression of emotions. In a real-life conversation, humans firstly instinctively perceive emotions from multi-source information, including the…

计算与语言 · 计算机科学 2022-03-31 Yunlong Liang , Fandong Meng , Ying Zhang , Jinan Xu , Yufeng Chen , Jie Zhou

Animating virtual avatars to make co-speech gestures facilitates various applications in human-machine interaction. The existing methods mainly rely on generative adversarial networks (GANs), which typically suffer from notorious mode…

计算机视觉与模式识别 · 计算机科学 2023-03-21 Lingting Zhu , Xian Liu , Xuanyu Liu , Rui Qian , Ziwei Liu , Lequan Yu

Audio-driven talking head generation necessitates seamless integration of audio and visual data amidst the challenges posed by diverse input portraits and intricate correlations between audio and facial motions. In response, we propose a…

计算机视觉与模式识别 · 计算机科学 2024-12-16 Ziqi Zhou , Weize Quan , Hailin Shi , Wei Li , Lili Wang , Dong-Ming Yan

Audio-driven emotional 3D face animation aims to generate emotionally expressive talking heads with synchronized lip movements. However, previous research has often overlooked the influence of diverse emotions on facial expressions or…

计算机视觉与模式识别 · 计算机科学 2024-08-06 Chang Liu , Qunfen Lin , Zijiao Zeng , Ye Pan

While accurate lip synchronization has been achieved for arbitrary-subject audio-driven talking face generation, the problem of how to efficiently drive the head pose remains. Previous methods rely on pre-estimated structural information…

计算机视觉与模式识别 · 计算机科学 2021-04-23 Hang Zhou , Yasheng Sun , Wayne Wu , Chen Change Loy , Xiaogang Wang , Ziwei Liu

In the era of advanced artificial intelligence and human-computer interaction, identifying emotions in spoken language is paramount. This research explores the integration of deep learning techniques in speech emotion recognition, offering…

声音 · 计算机科学 2023-10-20 Hanan Hamza , Fiza Gafoor , Fathima Sithara , Gayathri Anil , V. S. Anoop

Talking head generation is to synthesize a lip-synchronized talking head video by inputting an arbitrary face image and corresponding audio clips. Existing methods ignore not only the interaction and relationship of cross-modal information,…

计算机视觉与模式识别 · 计算机科学 2024-11-01 Sen Chen , Zhilei Liu , Jiaxing Liu , Longbiao Wang

Speech is a scalable and non-invasive biomarker for early mental health screening. However, widely used depression datasets like DAIC-WOZ exhibit strong coupling between linguistic sentiment and diagnostic labels, encouraging models to…

计算与语言 · 计算机科学 2026-01-05 Yuxin Li , Xiangyu Zhang , Yifei Li , Zhiwei Guo , Haoyang Zhang , Eng Siong Chng , Cuntai Guan

Image generation based on diffusion models has demonstrated impressive capability, motivating exploration into diverse and specialized applications. Owing to the importance of emotion in advertising, emotion-oriented image generation has…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Guoli Jia , Junyao Hu , Xinwei Long , Kai Tian , Kaiyan Zhang , KaiKai Zhao , Ning Ding , Bowen Zhou

Impressive progress has been made in audio-driven 3D facial animation recently, but synthesizing 3D talking-head with rich emotion is still unsolved. This is due to the lack of 3D generative models and available 3D emotional dataset with…

计算机视觉与模式识别 · 计算机科学 2021-04-27 Qianyun Wang , Zhenfeng Fan , Shihong Xia

In this paper, we present an Attention-based Identity Preserving Generative Adversarial Network (AIP-GAN) to overcome the identity leakage problem from a source image to a generated face image, an issue that is encountered in a…

计算机视觉与模式识别 · 计算机科学 2020-05-04 Kamran Ali , Charles E. Hughes

Generative models have attracted considerable attention for speech separation tasks, and among these, diffusion-based methods are being explored. Despite the notable success of diffusion techniques in generation tasks, their adaptation to…

音频与语音处理 · 电气工程与系统科学 2025-01-28 Jinwei Dong , Xinsheng Wang , Qirong Mao

Vision-guided speech generation aims to produce authentic speech from facial appearance or lip motions without relying on auditory signals, offering significant potential for applications such as dubbing in filmmaking and assisting…

计算机视觉与模式识别 · 计算机科学 2025-03-20 Jiaxin Ye , Hongming Shan

Person-generic audio-driven face generation is a challenging task in computer vision. Previous methods have achieved remarkable progress in audio-visual synchronization, but there is still a significant gap between current results and…

计算机视觉与模式识别 · 计算机科学 2024-08-09 Xiaozhong Ji , Chuming Lin , Zhonggan Ding , Ying Tai , Junwei Zhu , Xiaobin Hu , Donghao Luo , Yanhao Ge , Chengjie Wang

Emotion Recognition in Conversation (ERC) has attracted widespread attention in the natural language processing field due to its enormous potential for practical applications. Existing ERC methods face challenges in achieving generalization…

计算与语言 · 计算机科学 2023-09-20 Shanglin Lei , Xiaoping Wang , Guanting Dong , Jiang Li , Yingjian Liu