中文
相关论文

相关论文: EmoReg: Directional Latent Vector Modeling for Emo…

200 篇论文

Emotion serves as an essential component in daily human interactions. Existing human motion generation frameworks do not consider the impact of emotions, which reduces naturalness and limits their application in interactive tasks, such as…

计算机视觉与模式识别 · 计算机科学 2025-08-11 Chen Zhu , Buzhen Huang , Zijing Wu , Binghui Zuo , Yangang Wang

Word embedding models such as GloVe are widely used in natural language processing (NLP) research to convert words into vectors. Here, we provide a preliminary guide to probe latent emotions in text through GloVe word vectors. First, we…

计算与语言 · 计算机科学 2019-08-22 Zhengxuan Wu , Yueyi Jiang

We investigate hierarchical emotion distribution (ED) for achieving multi-level quantitative control of emotion rendering in text-to-speech synthesis (TTS). We introduce a novel multi-step hierarchical ED prediction module that quantifies…

声音 · 计算机科学 2025-07-08 Sho Inoue , Kun Zhou , Shuai Wang , Haizhou Li

Implementing fine-grained emotion control is crucial for emotion generation tasks because it enhances the expressive capability of the generative model, allowing it to accurately and comprehensively capture and express various nuanced…

计算机视觉与模式识别 · 计算机科学 2024-02-05 Guanwen Feng , Haoran Cheng , Yunan Li , Zhiyuan Ma , Chaoneng Li , Zhihao Qian , Qiguang Miao , Chi-Man Pun

Emotion Representation Mapping (ERM) has the goal to convert existing emotion ratings from one representation format into another one, e.g., mapping Valence-Arousal-Dominance annotations for words or sentences into Ekman's Basic Emotions…

计算与语言 · 计算机科学 2018-06-26 Sven Buechel , Udo Hahn

The domain of 3D talking head generation has witnessed significant progress in recent years. A notable challenge in this field consists in blending speech-related motions with expression dynamics, which is primarily caused by the lack of…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Federico Nocentini , Claudio Ferrari , Stefano Berretti

Emotion Recognition in Conversations (ERC) has been gaining increasing importance as conversational agents become more and more common. Recognizing emotions is key for effective communication, being a crucial component in the development of…

计算与语言 · 计算机科学 2023-06-06 Patrícia Pereira , Helena Moniz , Isabel Dias , Joao Paulo Carvalho

With the development of deep neural networks, automatic music composition has made great progress. Although emotional music can evoke listeners' different emotions and it is important for artistic expression, only few researches have…

声音 · 计算机科学 2021-12-17 Kaitong Zheng , Ruijie Meng , Chengshi Zheng , Xiaodong Li , Jinqiu Sang , Juanjuan Cai , Jie Wang

How much can we infer about an emotional voice solely from an expressive face? This intriguing question holds great potential for applications such as virtual character dubbing and aiding individuals with expressive language disorders.…

声音 · 计算机科学 2025-02-04 Jiaxin Ye , Boyuan Cao , Hongming Shan

The majority of current systems for end-to-end dialog generation focus on response quality without an explicit control over the affective content of the responses. In this paper, we present an affect-driven dialog system, which generates…

计算与语言 · 计算机科学 2019-04-08 Pierre Colombo , Wojciech Witon , Ashutosh Modi , James Kennedy , Mubbasir Kapadia

Emotional text-to-speech (TTS) technology has achieved significant progress in recent years; however, challenges remain owing to the inherent complexity of emotions and limitations of the available emotional speech datasets and models.…

声音 · 计算机科学 2025-04-18 Deok-Hyeon Cho , Hyung-Seok Oh , Seung-Bin Kim , Seong-Whan Lee

Diffusion models have shown promising results for a wide range of generative tasks with continuous data, such as image and audio synthesis. However, little progress has been made on using diffusion models to generate discrete symbolic music…

声音 · 计算机科学 2023-10-24 Jincheng Zhang , György Fazekas , Charalampos Saitis

Video grounding aims to localize the target moment in an untrimmed video corresponding to a given sentence query. Existing methods typically select the best prediction from a set of predefined proposals or directly regress the target span…

计算机视觉与模式识别 · 计算机科学 2024-01-01 Xiao Liang , Tao Shi , Yaoyuan Liang , Te Tao , Shao-Lun Huang

Recent advances in emotional voice conversion (EVC) have enabled the generation of expressive synthetic speech, raising new concerns in audio deepfake detection. Existing approaches treat speech as a homogeneous signal and largely overlook…

声音 · 计算机科学 2026-05-06 Vamshi Nallaguntla , Shruti Kshirsagar , Anderson R. Avila

Speech Emotion Conversion aims to modify the emotion expressed in input speech while preserving lexical content and speaker identity. Recently, generative modeling approaches have shown promising results in changing local acoustic…

音频与语音处理 · 电气工程与系统科学 2025-08-18 Navin Raj Prabhu , Danilo de Oliveira , Nale Lehmann-Willenbrock , Timo Gerkmann

This paper introduces a novel voice conversion (VC) model, guided by text instructions such as "articulate slowly with a deep tone" or "speak in a cheerful boyish voice". Unlike traditional methods that rely on reference utterances to…

音频与语音处理 · 电气工程与系统科学 2024-01-17 Chun-Yi Kuan , Chen An Li , Tsu-Yuan Hsu , Tse-Yang Lin , Ho-Lam Chung , Kai-Wei Chang , Shuo-yiin Chang , Hung-yi Lee

Existing gesture generation methods primarily focus on upper body gestures based on audio features, neglecting speech content, emotion, and locomotion. These limitations result in stiff, mechanical gestures that fail to convey the true…

声音 · 计算机科学 2026-03-10 Yongkang Cheng , Mingjiang Liang , Shaoli Huang , Gaoge Han , Jifeng Ning , Wei Liu

Diffusion models have shown great success in generating high-quality co-speech gestures for interactive humanoid robots or digital avatars from noisy input with the speech audio or text as conditions. However, they rarely focus on providing…

人机交互 · 计算机科学 2024-04-04 Zeyu Zhao , Nan Gao , Zhi Zeng , Guixuan Zhang , Jie Liu , Shuwu Zhang

Although voice conversion (VC) systems have shown a remarkable ability to transfer voice style, existing methods still have an inaccurate pitch and low speaker adaptation quality. To address these challenges, we introduce Diff-HierVC, a…

音频与语音处理 · 电气工程与系统科学 2023-11-09 Ha-Yeong Choi , Sang-Hoon Lee , Seong-Whan Lee

Over the recent years, various deep learning-based methods were proposed for extracting a fixed-dimensional embedding vector from speech signals. Although the deep learning-based embedding extraction methods have shown good performance in…

音频与语音处理 · 电气工程与系统科学 2021-12-08 Woo Hyun Kang , Jahangir Alam , Abderrahim Fathan