English
Related papers

Related papers: ED-TTS: Multi-Scale Emotion Modeling using Cross-D…

200 papers

Emotional text-to-speech synthesis (TTS) aims to generate realistic emotional speech from input text. However, quantitatively controlling multi-level emotion rendering remains challenging. In this paper, we propose a flow-matching based…

Sound · Computer Science 2025-06-24 Sho Inoue , Kun Zhou , Shuai Wang , Haizhou Li

Neural text-to-speech (TTS) approaches generally require a huge number of high quality speech data, which makes it difficult to obtain such a dataset with extra emotion labels. In this paper, we propose a novel approach for emotional TTS…

Audio and Speech Processing · Electrical Eng. & Systems 2021-01-19 Xiong Cai , Dongyang Dai , Zhiyong Wu , Xiang Li , Jingbei Li , Helen Meng

We often verbally express emotions in a multifaceted manner, they may vary in their intensities and may be expressed not just as a single but as a mixture of emotions. This wide spectrum of emotions is well-studied in the structural model…

Computation and Language · Computer Science 2024-06-28 Rendi Chevi , Alham Fikri Aji

There has been significant progress in emotional Text-To-Speech (TTS) synthesis technology in recent years. However, existing methods primarily focus on the synthesis of a limited number of emotion types and have achieved unsatisfactory…

Sound · Computer Science 2023-06-02 Haobin Tang , Xulong Zhang , Jianzong Wang , Ning Cheng , Jing Xiao

It remains a challenge to effectively control the emotion rendering in text-to-speech (TTS) synthesis. Prior studies have primarily focused on learning a global prosodic representation at the utterance level, which strongly correlates with…

Sound · Computer Science 2024-05-16 Sho Inoue , Kun Zhou , Shuai Wang , Haizhou Li

We investigate hierarchical emotion distribution (ED) for achieving multi-level quantitative control of emotion rendering in text-to-speech synthesis (TTS). We introduce a novel multi-step hierarchical ED prediction module that quantifies…

Sound · Computer Science 2025-07-08 Sho Inoue , Kun Zhou , Shuai Wang , Haizhou Li

Cross-speaker emotion transfer in speech synthesis relies on extracting speaker-independent emotion embeddings for accurate emotion modeling without retaining speaker traits. However, existing timbre compression methods fail to fully…

Sound · Computer Science 2025-10-20 Deok-Hyeon Cho , Hyung-Seok Oh , Seung-Bin Kim , Seong-Whan Lee

Cross-lingual emotional text-to-speech (TTS) aims to produce speech in one language that captures the emotion of a speaker from another language while maintaining the target voice's timbre. This process of cross-lingual emotional speech…

While the performance of cross-lingual TTS based on monolingual corpora has been significantly improved recently, generating cross-lingual speech still suffers from the foreign accent problem, leading to limited naturalness. Besides,…

Sound · Computer Science 2023-09-06 Tao Li , Chenxu Hu , Jian Cong , Xinfa Zhu , Jingbei Li , Qiao Tian , Yuping Wang , Lei Xie

The cross-speaker emotion transfer task in text-to-speech (TTS) synthesis particularly aims to synthesize speech for a target speaker with the emotion transferred from reference speech recorded by another (source) speaker. During the…

Sound · Computer Science 2022-04-11 Tao Li , Xinsheng Wang , Qicong Xie , Zhichao Wang , Lei Xie

In recent years, emotional Text-to-Speech (TTS) synthesis and emphasis-controllable speech synthesis have advanced significantly. However, their interaction remains underexplored. We propose Emphasis Meets Emotion TTS (EME-TTS), a novel…

Sound · Computer Science 2025-07-17 Haoxun Li , Leyuan Qu , Jiaxi Hu , Taihao Li

This paper proposes an effective emotion control method for an end-to-end text-to-speech (TTS) system. To flexibly control the distinct characteristic of a target emotion category, it is essential to determine embedding vectors representing…

Audio and Speech Processing · Electrical Eng. & Systems 2019-11-07 Se-Yun Um , Sangshin Oh , Kyungguen Byun , Inseon Jang , Chunghyun Ahn , Hong-Goo Kang

Expressive synthetic speech is essential for many human-computer interaction and audio broadcast scenarios, and thus synthesizing expressive speech has attracted much attention in recent years. Previous methods performed the expressive…

Sound · Computer Science 2022-01-19 Yi Lei , Shan Yang , Xinsheng Wang , Lei Xie

Current emotional Text-To-Speech (TTS) and style transfer methods rely on reference encoders to control global style or emotion vectors, but do not capture nuanced acoustic details of the reference speech. To this end, we propose a novel…

Sound · Computer Science 2025-10-03 Jianing Yang , Sheng Li , Takahiro Shinozaki , Yuki Saito , Hiroshi Saruwatari

Previous multilingual text-to-speech (TTS) approaches have considered leveraging monolingual speaker data to enable cross-lingual speech synthesis. However, such data-efficient approaches have ignored synthesizing emotional aspects of…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-01 Xinfa Zhu , Yi Lei , Tao Li , Yongmao Zhang , Hongbin Zhou , Heng Lu , Lei Xie

Recently, there has been an increasing interest in neural speech synthesis. While the deep neural network achieves the state-of-the-art result in text-to-speech (TTS) tasks, how to generate a more emotional and more expressive speech is…

Computation and Language · Computer Science 2021-06-24 Chenye Cui , Yi Ren , Jinglin Liu , Feiyang Chen , Rongjie Huang , Ming Lei , Zhou Zhao

In this paper, we propose to utilise diffusion models for data augmentation in speech emotion recognition (SER). In particular, we present an effective approach to utilise improved denoising diffusion probabilistic models (IDDPM) to…

Sound · Computer Science 2023-05-22 Ibrahim Malik , Siddique Latif , Raja Jurdak , Björn Schuller

Emotional text-to-speech synthesis (ETTS) has seen much progress in recent years. However, the generated voice is often not perceptually identifiable by its intended emotion category. To address this problem, we propose a new interactive…

Computation and Language · Computer Science 2021-06-15 Rui Liu , Berrak Sisman , Haizhou Li

While controllable Text-to-Speech (TTS) has achieved notable progress, most existing methods remain limited to inter-utterance-level control, making fine-grained intra-utterance expression challenging due to their reliance on non-public…

Sound · Computer Science 2026-05-19 Qifan Liang , Yuansen Liu , Ruixin Wei , Nan Lu , Junchuan Zhao , Ye Wang

Learning emotion embedding from reference audio is a straightforward approach for multi-emotion speech synthesis in encoder-decoder systems. But how to get better emotion embedding and how to inject it into TTS acoustic model more…

Sound · Computer Science 2022-01-31 Fengyu Yang , Jian Luan , Yujun Wang
‹ Prev 1 2 3 10 Next ›