English
Related papers

Related papers: EmoCat: Language-agnostic Emotional Voice Conversi…

200 papers

Audio-driven talking-head synthesis is a popular research topic for virtual human-related applications. However, the inflexibility and inefficiency of existing methods, which necessitate expensive end-to-end training to transfer emotions…

Sound · Computer Science 2023-10-13 Yuan Gan , Zongxin Yang , Xihang Yue , Lingyun Sun , Yi Yang

Humans can effortlessly modify various prosodic attributes, such as the placement of stress and the intensity of sentiment, to convey a specific emotion while maintaining consistent linguistic content. Motivated by this capability, we…

Sound · Computer Science 2023-12-29 Leyuan Qu , Wei Wang , Cornelius Weber , Pengcheng Yue , Taihao Li , Stefan Wermter

Speech emotion conversion is the task of modifying the perceived emotion of a speech utterance while preserving the lexical content and speaker identity. In this study, we cast the problem of emotion conversion as a spoken language…

Emotional voice conversion aims to convert the emotion of speech from one state to another while preserving the linguistic content and speaker identity. The prior studies on emotional voice conversion are mostly carried out under the…

Sound · Computer Science 2020-10-14 Kun Zhou , Berrak Sisman , Mingyang Zhang , Haizhou Li

We propose a nonparallel data-driven emotional speech conversion method. It enables the transfer of emotion-related characteristics of a speech signal while preserving the speaker's identity and linguistic content. Most existing approaches…

Machine Learning · Computer Science 2020-04-14 Jian Gao , Deep Chakraborty , Hamidou Tembine , Olaitan Olaleye

We present eCat, a novel end-to-end multispeaker model capable of: a) generating long-context speech with expressive and contextually appropriate prosody, and b) performing fine-grained prosody transfer between any pair of seen speakers.…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-21 Ammar Abbas , Sri Karlapati , Bastian Schnell , Penny Karanasou , Marcel Granero Moya , Amith Nagaraj , Ayman Boustati , Nicole Peinelt , Alexis Moinet , Thomas Drugman

Speech Emotion Recognition is a crucial area of research in human-computer interaction. While significant work has been done in this field, many state-of-the-art networks struggle to accurately recognize emotions in speech when the data is…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-23 Rashedul Hasan , Meher Nigar , Nursadul Mamun , Sayan Paul

We propose EmoLat, a novel emotion latent space that enables fine-grained, text-driven image sentiment transfer by modeling cross-modal correlations between textual semantics and visual emotion features. Within EmoLat, an emotion semantic…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Jing Zhang , Bingjie Fan , Jixiang Zhu , Zhe Wang

Many datasets have been designed to further the development of fake audio detection, such as datasets of the ASVspoof and ADD challenges. However, these datasets do not consider a situation that the emotion of the audio has been changed…

Sound · Computer Science 2024-07-25 Yan Zhao , Jiangyan Yi , Jianhua Tao , Chenglong Wang , Xiaohui Zhang , Yongfeng Dong

Emotional voice conversion (EVC) seeks to convert the emotional state of an utterance while preserving the linguistic content and speaker identity. In EVC, emotions are usually treated as discrete categories overlooking the fact that speech…

Sound · Computer Science 2022-07-19 Kun Zhou , Berrak Sisman , Rajib Rana , Björn W. Schuller , Haizhou Li

Recent advances in zero-shot voice conversion have exhibited potential in emotion control, yet the performance is suboptimal or inconsistent due to their limited expressive capacity. We propose Emotion-Aware Prefix for explicit emotion…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-11 Haoyuan Yang , Mu Yang , Jiamin Xie , Szu-Jui Chen , John H. L. Hansen

Emotional talking head synthesis aims to generate talking portrait videos with vivid expressions. Existing methods still exhibit limitations in control flexibility, motion naturalness, and expression quality. Moreover, currently available…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Yiguo Jiang , Xiaodong Cun , Yong Zhang , Yudian Zheng , Fan Tang , Chi-Man Pun

Speech emotion conversion is the task of converting the expressed emotion of a spoken utterance to a target emotion while preserving the lexical content and speaker identity. While most existing works in speech emotion conversion rely on…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-09 Navin Raj Prabhu , Bunlong Lay , Simon Welker , Nale Lehmann-Willenbrock , Timo Gerkmann

Human speech goes beyond the mere transfer of information; it is a profound exchange of emotions and a connection between individuals. While Text-to-Speech (TTS) models have made huge progress, they still face challenges in controlling the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-14 Guanrou Yang , Chen Yang , Qian Chen , Ziyang Ma , Wenxi Chen , Wen Wang , Tianrui Wang , Yifan Yang , Zhikang Niu , Wenrui Liu , Fan Yu , Zhihao Du , Zhifu Gao , ShiLiang Zhang , Xie Chen

Emotional voice conversion aims to transform emotional prosody in speech while preserving the linguistic content and speaker identity. Prior studies show that it is possible to disentangle emotional prosody using an encoder-decoder network…

Sound · Computer Science 2021-02-12 Kun Zhou , Berrak Sisman , Rui Liu , Haizhou Li

The capability of generating speech with specific type of emotion is desired for many applications of human-computer interaction. Cross-speaker emotion transfer is a common approach to generating emotional speech when speech with emotion…

Audio and Speech Processing · Electrical Eng. & Systems 2023-01-05 Guangyan Zhang , Ying Qin , Wenjie Zhang , Jialun Wu , Mei Li , Yutao Gai , Feijun Jiang , Tan Lee

Emotional Voice Conversion (EVC) aims to convert the emotional style of a source speech signal to a target style while preserving its content and speaker identity information. Previous emotional conversion studies do not disentangle…

Sound · Computer Science 2021-07-20 Xiangheng He , Junjie Chen , Georgios Rizos , Björn W. Schuller

We propose a novel transfer learning method for speech emotion recognition allowing us to obtain promising results when only few training data is available. With as low as 125 examples per emotion class, we were able to reach a higher…

Machine Learning · Computer Science 2020-11-12 Jonathan Boigne , Biman Liyanage , Ted Östrem

Emotional voice conversion (EVC) is one way to generate expressive synthetic speech. Previous approaches mainly focused on modeling one-to-one mapping, i.e., conversion from one emotional state to another emotional state, with Mel-cepstral…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-09 Songxiang Liu , Yuewen Cao , Helen Meng

Speech emotion recognition plays an important role in building more intelligent and human-like agents. Due to the difficulty of collecting speech emotional data, an increasingly popular solution is leveraging a related and rich source…

Machine Learning · Computer Science 2019-02-15 Hao Zhou , Ke Chen
‹ Prev 1 2 3 10 Next ›