中文
相关论文

相关论文: MoEE: Mixture of Emotion Experts for Audio-Driven …

200 篇论文

In this work, we tackle the challenge of enhancing the realism and expressiveness in talking head video generation by focusing on the dynamic and nuanced relationship between audio cues and facial movements. We identify the limitations of…

计算机视觉与模式识别 · 计算机科学 2024-08-09 Linrui Tian , Qi Wang , Bang Zhang , Liefeng Bo

Emotion Recognition in Conversations (ERC) presents unique challenges, requiring models to capture the temporal flow of multi-turn dialogues and to effectively integrate cues from multiple modalities. We propose Mixture of Speech-Text…

计算与语言 · 计算机科学 2026-02-27 Soumya Dutta , Smruthi Balaji , Sriram Ganapathy

Speech-driven 3D facial animation has garnered lots of attention thanks to its broad range of applications. Despite recent advancements in achieving realistic lip motion, current methods fail to capture the nuanced emotional undertones…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Jisoo Kim , Jungbin Cho , Joonho Park , Soonmin Hwang , Da Eun Kim , Geon Kim , Youngjae Yu

It is in high demand to generate facial animation with high realism, but it remains a challenging task. Existing approaches of speech-driven facial animation can produce satisfactory mouth movement and lip synchronization, but show weakness…

计算机视觉与模式识别 · 计算机科学 2025-05-08 Yutong Chen , Junhong Zhao , Wei-Qiang Zhang

Head avatars animated by visual signals have gained popularity, particularly in cross-driving synthesis where the driver differs from the animated character, a challenging but highly practical approach. The recently presented MegaPortraits…

计算机视觉与模式识别 · 计算机科学 2024-05-01 Nikita Drobyshev , Antoni Bigata Casademunt , Konstantinos Vougioukas , Zoe Landgraf , Stavros Petridis , Maja Pantic

Recent advances in video diffusion models have unlocked new potential for realistic audio-driven talking video generation. However, achieving seamless audio-lip synchronization, maintaining long-term identity consistency, and producing…

计算机视觉与模式识别 · 计算机科学 2024-12-06 Longtao Zheng , Yifan Zhang , Hanzhong Guo , Jiachun Pan , Zhenxiong Tan , Jiahao Lu , Chuanxin Tang , Bo An , Shuicheng Yan

Diffusion models have revolutionized the field of talking head generation, yet still face challenges in expressiveness, controllability, and stability in long-time generation. In this research, we propose an EmotiveTalk framework to address…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Haotian Wang , Yuzhe Weng , Yueyan Li , Zilu Guo , Jun Du , Shutong Niu , Jiefeng Ma , Shan He , Xiaoyan Wu , Qiming Hu , Bing Yin , Cong Liu , Qingfeng Liu

Audio-driven talking head generation necessitates seamless integration of audio and visual data amidst the challenges posed by diverse input portraits and intricate correlations between audio and facial motions. In response, we propose a…

计算机视觉与模式识别 · 计算机科学 2024-12-16 Ziqi Zhou , Weize Quan , Hailin Shi , Wei Li , Lili Wang , Dong-Ming Yan

To be widely adopted, 3D facial avatars must be animated easily, realistically, and directly from speech signals. While the best recent methods generate 3D animations that are synchronized with the input audio, they largely ignore the…

计算机视觉与模式识别 · 计算机科学 2023-09-27 Radek Daněček , Kiran Chhatre , Shashank Tripathi , Yandong Wen , Michael J. Black , Timo Bolkart

Audio-driven talking-head generation is a crucial and useful technology for virtual human interaction and film-making. While recent advances have focused on improving image fidelity and lip synchronization, generating accurate emotional…

计算机视觉与模式识别 · 计算机科学 2025-05-05 Wenqing Wang , Yun Fu

3D facial animation has attracted considerable attention due to its extensive applications in the multimedia field. Audio-driven 3D facial animation has been widely explored with promising results. However, multi-modal 3D facial animation,…

计算机视觉与模式识别 · 计算机科学 2024-10-11 Sijing Wu , Yunhao Li , Yichao Yan , Huiyu Duan , Ziwei Liu , Guangtao Zhai

Enabling digital humans to express rich emotions has significant applications in dialogue systems, gaming, and other interactive scenarios. While recent advances in talking head synthesis have achieved impressive results in lip…

人工智能 · 计算机科学 2025-10-21 Haidong Xu , Meishan Zhang , Hao Ju , Zhedong Zheng , Erik Cambria , Min Zhang , Hao Fei

Advances in multimodal models have greatly improved how interactions relevant to various tasks are modeled. Today's multimodal models mainly focus on the correspondence between images and text, using this for tasks like image-text matching.…

计算与语言 · 计算机科学 2024-09-27 Haofei Yu , Zhengyang Qi , Lawrence Jang , Ruslan Salakhutdinov , Louis-Philippe Morency , Paul Pu Liang

Emotional voice conversion (EVC) traditionally targets the transformation of spoken utterances from one emotional state to another, with previous research mainly focusing on discrete emotion categories. This paper departs from the norm by…

音频与语音处理 · 电气工程与系统科学 2023-09-19 Kun Zhou , Berrak Sisman , Carlos Busso , Bin Ma , Haizhou Li

Several works have developed end-to-end pipelines for generating lip-synced talking faces with various real-world applications, such as teaching and language translation in videos. However, these prior works fail to create realistic-looking…

计算机视觉与模式识别 · 计算机科学 2023-03-28 Sahil Goyal , Shagun Uppal , Sarthak Bhagat , Yi Yu , Yifang Yin , Rajiv Ratn Shah

The task of audio-driven portrait animation involves generating a talking head video using an identity image and an audio track of speech. While many existing approaches focus on lip synchronization and video quality, few tackle the…

计算机视觉与模式识别 · 计算机科学 2024-09-12 Jian Zhang , Weijian Mai , Zhijun Zhang

Multimodal emotion recognition (MER) is crucial for human-computer interaction, yet real-world challenges like dynamic modality incompleteness and asynchrony severely limit its robustness. Existing methods often assume consistently complete…

人机交互 · 计算机科学 2025-08-19 Yitong Zhu , Lei Han , Guanxuan Jiang , PengYuan Zhou , Yuyang Wang

In this paper, we propose a novel audio-driven talking head method capable of simultaneously generating highly expressive facial expressions and hand gestures. Unlike existing methods that focus on generating full-body or half-body poses,…

计算机视觉与模式识别 · 计算机科学 2025-01-22 Linrui Tian , Siqi Hu , Qi Wang , Bang Zhang , Liefeng Bo

Understanding emotions accurately is essential for fields like human-computer interaction. Due to the complexity of emotions and their multi-modal nature (e.g., emotions are influenced by facial expressions and audio), researchers have…

计算机视觉与模式识别 · 计算机科学 2025-01-17 Qize Yang , Detao Bai , Yi-Xing Peng , Xihan Wei

Talking face generation is a novel and challenging generation task, aiming at synthesizing a vivid speaking-face video given a specific audio. To fulfill emotion-controllable talking face generation, current methods need to overcome two…

计算机视觉与模式识别 · 计算机科学 2025-08-21 Ziqi Zhang , Cheng Deng
‹ 上一页 1 2 3 10 下一页 ›