English
Related papers

Related papers: CoCoEmo: Composable and Controllable Human-Like Em…

200 papers

Cross-lingual emotional text-to-speech (TTS) aims to produce speech in one language that captures the emotion of a speaker from another language while maintaining the target voice's timbre. This process of cross-lingual emotional speech…

Current emotional text-to-speech (TTS) models predominantly conduct supervised training to learn the conversion from text and desired emotion to its emotional speech, focusing on a single emotion per text-speech pair. These models only…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-17 Xiaoxue Gao , Chen Zhang , Yiming Chen , Huayun Zhang , Nancy F. Chen

While current emotional Text-to-Speech (TTS) models have successfully controlled verbal prosody, they often ignore non-verbal vocalizations (NVs), which are essential for authentic human emotion. Although some non-verbal datasets have…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-26 Wangzixi Zhou , Bagus Tris Atmaja , Sakriani Sakti

Emotional Text-to-Speech (E-TTS) synthesis has garnered significant attention in recent years due to its potential to revolutionize human-computer interaction. However, current E-TTS approaches often struggle to capture the intricacies of…

Computation and Language · Computer Science 2025-02-20 Zhi-Qi Cheng , Xiang Li , Jun-Yan He , Junyao Chen , Xiaomao Fan , Xiaojiang Peng , Alexander G. Hauptmann

Although current neural text-to-speech (TTS) models are able to generate high-quality speech, intensity controllable emotional TTS is still a challenging task. Most existing methods need external optimizations for intensity calculation,…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-17 Yiwei Guo , Chenpeng Du , Xie Chen , Kai Yu

Advances in text-to-speech (TTS) technology have significantly improved the quality of generated speech, closely matching the timbre and intonation of the target speaker. However, due to the inherent complexity of human emotional…

Sound · Computer Science 2024-12-13 Weizhen Bian , Yubo Zhou , Kaitai Zhang , Xiaohan Gu

This paper proposes an effective emotional text-to-speech (TTS) system with a pre-trained language model (LM)-based emotion prediction method. Unlike conventional systems that require auxiliary inputs such as manually defined emotion…

While Text-to-Speech (TTS) systems enable emotional control via natural-language instructions, expressiveness, naturalness, and speech quality degrade when the target emotion conflicts with the textual semantics. We propose a Cross-modal…

Computation and Language · Computer Science 2026-05-20 Yizhou Peng , Yukun Ma , Chong Zhang , Yi-Wen Chao , Chongjia Ni , Bin Ma , Eng Siong Chng

We often verbally express emotions in a multifaceted manner, they may vary in their intensities and may be expressed not just as a single but as a mixture of emotions. This wide spectrum of emotions is well-studied in the structural model…

Computation and Language · Computer Science 2024-06-28 Rendi Chevi , Alham Fikri Aji

Purpose: Emotion is a fundamental component of human communication, shaping understanding, trust, and engagement across domains such as education, healthcare, and mental health. While large language models (LLMs) exhibit strong reasoning…

Computation and Language · Computer Science 2025-10-15 Yurui Dong , Luozhijie Jin , Yao Yang , Bingjie Lu , Jiaxi Yang , Zhi Liu

Previous multilingual text-to-speech (TTS) approaches have considered leveraging monolingual speaker data to enable cross-lingual speech synthesis. However, such data-efficient approaches have ignored synthesizing emotional aspects of…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-01 Xinfa Zhu , Yi Lei , Tao Li , Yongmao Zhang , Hongbin Zhou , Heng Lu , Lei Xie

Neural text-to-speech (TTS) approaches generally require a huge number of high quality speech data, which makes it difficult to obtain such a dataset with extra emotion labels. In this paper, we propose a novel approach for emotional TTS…

Audio and Speech Processing · Electrical Eng. & Systems 2021-01-19 Xiong Cai , Dongyang Dai , Zhiyong Wu , Xiang Li , Jingbei Li , Helen Meng

Large Language Models (LLMs) demonstrate increasing conversational fluency, yet instilling them with nuanced, human-like emotional expression remains a significant challenge. Current alignment techniques often address surface-level output…

Computation and Language · Computer Science 2025-11-25 Niranjan Chebrolu , Gerard Christopher Yeo , Kokil Jaidka

Deploying LLMs in real-world applications requires controllable output that satisfies multiple desiderata at the same time. While existing work extensively addresses LLM steering for a single behavior, \textit{compositional steering} --…

Computation and Language · Computer Science 2026-04-21 Gorjan Radevski , Kiril Gashteovski , Giwon Hong , Carolin Lawrence , Goran Glavaš

Emotional text-to-speech (E-TTS) is central to creating natural and trustworthy human-computer interaction. Existing systems typically rely on sentence-level control through predefined labels, reference audio, or natural language prompts.…

Computation and Language · Computer Science 2025-09-26 Sirui Wang , Andong Chen , Tiejun Zhao

Although current Text-To-Speech (TTS) models are able to generate high-quality speech samples, there are still challenges in developing emotion intensity controllable TTS. Most existing TTS models achieve emotion intensity control by…

Sound · Computer Science 2024-05-28 Haoxiang Shi , Jianzong Wang , Xulong Zhang , Ning Cheng , Jun Yu , Jing Xiao

Turn-taking is a fundamental aspect of human communication where speakers convey their intention to either hold, or yield, their turn through prosodic cues. Using the recently proposed Voice Activity Projection model, we propose an…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-30 Erik Ekstedt , Siyang Wang , Éva Székely , Joakim Gustafson , Gabriel Skantze

Human speech goes beyond the mere transfer of information; it is a profound exchange of emotions and a connection between individuals. While Text-to-Speech (TTS) models have made huge progress, they still face challenges in controlling the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-14 Guanrou Yang , Chen Yang , Qian Chen , Ziyang Ma , Wenxi Chen , Wen Wang , Tianrui Wang , Yifan Yang , Zhikang Niu , Wenrui Liu , Fan Yu , Zhihao Du , Zhifu Gao , ShiLiang Zhang , Xie Chen

Recent advancements in speech synthesis have enabled large language model (LLM)-based systems to perform zero-shot generation with controllable content, timbre, speaker identity, and emotion through input prompts. As a result, these models…

Expressive text-to-speech (TTS) aims to synthesize speeches with human-like tones, moods, or even artistic attributes. Recent advancements in expressive TTS empower users with the ability to directly control synthesis style through natural…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-03 Hanglei Zhang , Yiwei Guo , Sen Liu , Xie Chen , Kai Yu