English
Related papers

Related papers: EmoCAST: Emotional Talking Portrait via Emotive Te…

200 papers

Emotional support is a crucial ability for many conversation scenarios, including social interactions, mental health support, and customer service chats. Following reasonable procedures and using various support skills can help to…

Computation and Language · Computer Science 2021-06-03 Siyang Liu , Chujie Zheng , Orianna Demasi , Sahand Sabour , Yu Li , Zhou Yu , Yong Jiang , Minlie Huang

Although automatically animating audio-driven talking heads has recently received growing interest, previous efforts have mainly concentrated on achieving lip synchronization with the audio, neglecting two crucial elements for generating…

Computer Vision and Pattern Recognition · Computer Science 2024-03-13 Shuai Tan , Bin Ji , Ye Pan

Humans can effortlessly modify various prosodic attributes, such as the placement of stress and the intensity of sentiment, to convey a specific emotion while maintaining consistent linguistic content. Motivated by this capability, we…

Sound · Computer Science 2023-12-29 Leyuan Qu , Wei Wang , Cornelius Weber , Pengcheng Yue , Taihao Li , Stefan Wermter

Emotion is essential in spoken communication, yet most existing frameworks in speech emotion modeling rely on predefined categories or low-dimensional continuous attributes, which offer limited expressive capacity. Recent advances in speech…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-07 Tianhua Qi , Wenming Zheng , Björn W. Schuller , Zhaojie Luo , Haizhou Li

Although significant progress has been made to audio-driven talking face generation, existing methods either neglect facial emotion or cannot be applied to arbitrary subjects. In this paper, we propose the Emotion-Aware Motion Model (EAMM)…

Computer Vision and Pattern Recognition · Computer Science 2022-09-26 Xinya Ji , Hang Zhou , Kaisiyuan Wang , Qianyi Wu , Wayne Wu , Feng Xu , Xun Cao

Story generation aims to produce image sequences that depict coherent narratives while maintaining subject consistency across frames. Although existing methods have excelled in producing coherent and expressive stories, they remain largely…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Jingyuan Yang , Rucong Chen , Weibin Luo , Hui Huang

Despite the rapid progress in image generation, emotional image editing remains under-explored. The semantics, context, and structure of an image can evoke emotional responses, making emotional image editing techniques valuable for various…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Qing Lin , Jingfeng Zhang , Yew-Soon Ong , Mengmi Zhang

We propose EmoLat, a novel emotion latent space that enables fine-grained, text-driven image sentiment transfer by modeling cross-modal correlations between textual semantics and visual emotion features. Within EmoLat, an emotion semantic…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Jing Zhang , Bingjie Fan , Jixiang Zhu , Zhe Wang

Emotions play a central role in human communication, shaping trust, engagement, and social interaction. As artificial intelligence systems powered by large language models become increasingly integrated into everyday life, enabling them to…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-11 Soumya Dutta

Conversational Speech Synthesis (CSS) aims to accurately express an utterance with the appropriate prosody and emotional inflection within a conversational setting. While recognising the significance of CSS task, the prior studies have not…

Computation and Language · Computer Science 2023-12-20 Rui Liu , Yifan Hu , Yi Ren , Xiang Yin , Haizhou Li

Humans can perceive speakers' characteristics (e.g., identity, gender, personality and emotion) by their appearance, which are generally aligned to their voice style. Recently, vision-driven Text-to-speech (TTS) scholars grounded their…

Sound · Computer Science 2025-04-17 Tian-Hao Zhang , Jiawei Zhang , Jun Wang , Xinyuan Qian , Xu-Cheng Yin

The generation of stylistic 3D facial animations driven by speech presents a significant challenge as it requires learning a many-to-many mapping between speech, style, and the corresponding natural facial motion. However, existing methods…

Computer Vision and Pattern Recognition · Computer Science 2024-05-15 Zhiyao Sun , Tian Lv , Sheng Ye , Matthieu Lin , Jenny Sheng , Yu-Hui Wen , Minjing Yu , Yong-Jin Liu

Speech-driven 3D facial animation aims to generate realistic and expressive facial motions directly from audio. While recent methods achieve high-quality lip synchronization, they often rely on discrete emotion categories, limiting…

Multimedia · Computer Science 2026-01-16 Diqiong Jiang , Kai Zhu , Dan Song , Jian Chang , Chenglizhao Chen , Zhenyu Wu

The generation of talking avatars has achieved significant advancements in precise audio synchronization. However, crafting lifelike talking head videos requires capturing a broad spectrum of emotions and subtle facial expressions. Current…

Computer Vision and Pattern Recognition · Computer Science 2025-01-10 Huaize Liu , Wenzhang Sun , Donglin Di , Shibo Sun , Jiahui Yang , Changqing Zou , Hujun Bao

Emotional voice conversion (EVC) aims to change the emotional state of an utterance while preserving the linguistic content and speaker identity. In this paper, we propose a novel 2-stage training strategy for sequence-to-sequence emotional…

Computation and Language · Computer Science 2021-06-10 Kun Zhou , Berrak Sisman , Haizhou Li

Speech-driven 3D face animation aims to generate realistic facial expressions that match the speech content and emotion. However, existing methods often neglect emotional facial expressions or fail to disentangle them from speech content.…

Computer Vision and Pattern Recognition · Computer Science 2023-08-28 Ziqiao Peng , Haoyu Wu , Zhenbo Song , Hao Xu , Xiangyu Zhu , Jun He , Hongyan Liu , Zhaoxin Fan

Expressive synthetic speech is essential for many human-computer interaction and audio broadcast scenarios, and thus synthesizing expressive speech has attracted much attention in recent years. Previous methods performed the expressive…

Sound · Computer Science 2022-01-19 Yi Lei , Shan Yang , Xinsheng Wang , Lei Xie

Audio-driven portrait animation has made significant advances with diffusion-based models, improving video quality and lipsync accuracy. However, the increasing complexity of these models has led to inefficiencies in training and inference,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-17 Xuyang Cao , Guoxin Wang , Sheng Shi , Jun Zhao , Yang Yao , Jintao Fei , Minyu Gao , Pei Xie

Impressive progress has been made in audio-driven 3D facial animation recently, but synthesizing 3D talking-head with rich emotion is still unsolved. This is due to the lack of 3D generative models and available 3D emotional dataset with…

Computer Vision and Pattern Recognition · Computer Science 2021-04-27 Qianyun Wang , Zhenfeng Fan , Shihong Xia

Talking face generation is a novel and challenging generation task, aiming at synthesizing a vivid speaking-face video given a specific audio. To fulfill emotion-controllable talking face generation, current methods need to overcome two…

Computer Vision and Pattern Recognition · Computer Science 2025-08-21 Ziqi Zhang , Cheng Deng