中文
相关论文

相关论文: SEDTalker: Emotion-Aware 3D Facial Animation Using…

200 篇论文

End-to-End Neural Diarization with Encoder-Decoder based Attractor (EEND-EDA) is an end-to-end neural model for automatic speaker segmentation and labeling. It achieves the capability to handle flexible number of speakers by estimating the…

音频与语音处理 · 电气工程与系统科学 2024-03-22 PeiYing Lee , HauYun Guo , Berlin Chen

In filmmaking, directors typically allow actors to perform freely based on the script before providing specific guidance on how to present key actions. AI-generated content faces similar requirements, where users not only need automatic…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Zheng Qin , Ruobing Zheng , Yabing Wang , Tianqi Li , Zixin Zhu , Sanping Zhou , Ming Yang , Le Wang

Emotion recognition in conversations (ERC) is vital to the advancements of conversational AI and its applications. Therefore, the development of an automated ERC model using the concepts of machine learning (ML) would be beneficial.…

计算与语言 · 计算机科学 2023-06-06 Amitabha Dey , Shan Suthaharan

Speech-driven facial animation requires accurate correspondence between acoustic signals and facial motion, especially for articulation-related mouth movements. However, directly mapping speech audio to facial coefficients often overlooks…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Kai Zheng , Zejian Kang , Rui Mao , Hongyuan Zou , Yuanchen Fei , Xuanyang Xu , Xiangru Huang

Human conversation involves continuous exchanges of speech and nonverbal cues such as head nods, gaze shifts, and facial expressions that convey attention and emotion. Modeling these bidirectional dynamics in 3D is essential for building…

计算机视觉与模式识别 · 计算机科学 2026-02-27 Junjie Chen , Fei Wang , Zhihao Huang , Qing Zhou , Kun Li , Dan Guo , Linfeng Zhang , Xun Yang

Speech-Driven Facial Animation (SDFA) has gained significant attention due to its applications in movies, video games, and virtual reality. However, most existing models are trained on single-language data, limiting their effectiveness in…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Federico Nocentini , Kwanggyoon Seo , Qingju Liu , Claudio Ferrari , Stefano Berretti , David Ferman , Hyeongwoo Kim , Pablo Garrido , Akin Caliskan

Automated emotion detection in speech is a challenging task due to the complex interdependence between words and the manner in which they are spoken. It is made more difficult by the available datasets; their small size and incompatible…

音频与语音处理 · 电气工程与系统科学 2020-11-16 Amith Ananthram , Kailash Karthik Saravanakumar , Jessica Huynh , Homayoon Beigi

Human emotional expression is inherently dynamic, complex, and fluid, characterized by smooth transitions in intensity throughout verbal communication. However, the modeling of such intensity fluctuations has been largely overlooked by…

声音 · 计算机科学 2024-10-01 Jingyi Xu , Hieu Le , Zhixin Shu , Yang Wang , Yi-Hsuan Tsai , Dimitris Samaras

Most of the existing audio-driven 3D facial animation methods suffered from the lack of detailed facial expression and head pose, resulting in unsatisfactory experience of human-robot interaction. In this paper, a novel pose-controllable 3D…

计算机视觉与模式识别 · 计算机科学 2023-02-27 Bin Liu , Xiaolin Wei , Bo Li , Junjie Cao , Yu-Kun Lai

Generating realistic talking faces is a complex and widely discussed task with numerous applications. In this paper, we present DiffTalker, a novel model designed to generate lifelike talking faces through audio and landmark co-driving.…

计算机视觉与模式识别 · 计算机科学 2023-09-15 Zipeng Qi , Xulong Zhang , Ning Cheng , Jing Xiao , Jianzong Wang

Emotional talking-head generation has emerged as a pivotal research area at the intersection of computer vision and multimodal artificial intelligence, with its core value lying in enhancing human-computer interaction through immersive and…

计算机视觉与模式识别 · 计算机科学 2025-07-18 Hanlei Shi , Leyuan Qu , Yu Liu , Di Gao , Yuhua Zheng , Taihao Li

Speech-driven facial animation methods usually contain two main classes, 3D and 2D talking face, both of which attract considerable research attention in recent years. However, to the best of our knowledge, the research on 3D talking face…

计算机视觉与模式识别 · 计算机科学 2024-04-22 Yixiang Zhuang , Baoping Cheng , Yao Cheng , Yuntao Jin , Renshuai Liu , Chengyang Li , Xuan Cheng , Jing Liao , Juncong Lin

The objective of stylized speech-driven facial animation is to create animations that encapsulate specific emotional expressions. Existing methods often depend on pre-established emotional labels or facial expression templates, which may…

计算机视觉与模式识别 · 计算机科学 2023-09-12 Yicheng Zhong , Huawei Wei , Peiji Yang , Zhisheng Wang

Recently, 2D speaking avatars have increasingly participated in everyday scenarios due to the fast development of facial animation techniques. However, most existing works neglect the explicit control of human bodies. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Jiazhi Guan , Quanwei Yang , Kaisiyuan Wang , Hang Zhou , Shengyi He , Zhiliang Xu , Haocheng Feng , Errui Ding , Jingdong Wang , Hongtao Xie , Youjian Zhao , Ziwei Liu

Speech Emotion Recognition (SER) typically relies on utterance-level solutions. However, emotions conveyed through speech should be considered as discrete speech events with definite temporal boundaries, rather than attributes of the entire…

计算与语言 · 计算机科学 2023-10-23 Yingzhi Wang , Mirco Ravanelli , Alya Yacoubi

Individuals have unique facial expression and head pose styles that reflect their personalized speaking styles. Existing one-shot talking head methods cannot capture such personalized characteristics and therefore fail to produce diverse…

计算机视觉与模式识别 · 计算机科学 2024-09-17 Suzhen Wang , Yifeng Ma , Yu Ding , Zhipeng Hu , Changjie Fan , Tangjie Lv , Zhidong Deng , Xin Yu

Emotional talking head generation has attracted growing attention. Previous methods, which are mainly GAN-based, still struggle to consistently produce satisfactory results across diverse emotions and cannot conveniently specify…

计算机视觉与模式识别 · 计算机科学 2024-08-13 Yifeng Ma , Shiwei Zhang , Jiayu Wang , Xiang Wang , Yingya Zhang , Zhidong Deng

Audio-driven single-image talking portrait generation plays a crucial role in virtual reality, digital human creation, and filmmaking. Existing approaches are generally categorized into keypoint-based and image-based methods. Keypoint-based…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Chaolong Yang , Kai Yao , Yuyao Yan , Chenru Jiang , Weiguang Zhao , Jie Sun , Guangliang Cheng , Yifei Zhang , Bin Dong , Kaizhu Huang

Although significant progress has been made to audio-driven talking face generation, existing methods either neglect facial emotion or cannot be applied to arbitrary subjects. In this paper, we propose the Emotion-Aware Motion Model (EAMM)…

计算机视觉与模式识别 · 计算机科学 2022-09-26 Xinya Ji , Hang Zhou , Kaisiyuan Wang , Qianyi Wu , Wayne Wu , Feng Xu , Xun Cao

In this work, we present a lightweight and privacy-preserving Multimodal Emotion Recognition (MER) framework designed for deployment on edge devices. To demonstrate framework's versatility, our implementation uses three modalities - speech,…