中文
相关论文

相关论文: EmoSphere++: Emotion-Controllable Zero-Shot Text-t…

200 篇论文

The human voice conveys not just words but also emotional states and individuality. Emotional voice conversion (EVC) modifies emotional expressions while preserving linguistic content and speaker identity, improving applications like…

音频与语音处理 · 电气工程与系统科学 2025-09-29 Hsing-Hang Chou , Yun-Shao Lin , Ching-Chin Sung , Yu Tsao , Chi-Chun Lee

While Text-to-Speech (TTS) systems enable emotional control via natural-language instructions, expressiveness, naturalness, and speech quality degrade when the target emotion conflicts with the textual semantics. We propose a Cross-modal…

计算与语言 · 计算机科学 2026-05-20 Yizhou Peng , Yukun Ma , Chong Zhang , Yi-Wen Chao , Chongjia Ni , Bin Ma , Eng Siong Chng

Speech Emotion Recognition is a crucial area of research in human-computer interaction. While significant work has been done in this field, many state-of-the-art networks struggle to accurately recognize emotions in speech when the data is…

音频与语音处理 · 电气工程与系统科学 2025-01-23 Rashedul Hasan , Meher Nigar , Nursadul Mamun , Sayan Paul

Data augmentation via voice conversion (VC) has been successfully applied to low-resource expressive text-to-speech (TTS) when only neutral data for the target speaker are available. Although the quality of VC is crucial for this approach,…

音频与语音处理 · 电气工程与系统科学 2022-07-06 Ryo Terashima , Ryuichi Yamamoto , Eunwoo Song , Yuma Shirahata , Hyun-Wook Yoon , Jae-Min Kim , Kentaro Tachibana

Recently, large language model (LLM) based text-to-speech (TTS) systems have gradually become the mainstream in the industry due to their high naturalness and powerful zero-shot voice cloning capabilities.Here, we introduce the IndexTTS…

声音 · 计算机科学 2025-02-11 Wei Deng , Siyi Zhou , Jingchen Shu , Jinchao Wang , Lu Wang

Accent is an integral part of society, reflecting multiculturalism and shaping how individuals express identity. The majority of English speakers are non-native (L2) speakers, yet current Text-To-Speech (TTS) systems primarily model…

计算与语言 · 计算机科学 2026-03-10 Thanathai Lertpetchpun , Thanapat Trachu , Jihwan Lee , Tiantian Feng , Dani Byrd , Shrikanth Narayanan

Effective speech emotional representations play a key role in Speech Emotion Recognition (SER) and Emotional Text-To-Speech (TTS) tasks. However, emotional speech samples are more difficult and expensive to acquire compared with Neutral…

音频与语音处理 · 电气工程与系统科学 2023-06-12 Shijun Wang , Jón Guðnason , Damian Borth

We propose a Text-to-Speech method to create an unseen expressive style using one utterance of expressive speech of around one second. Specifically, we enhance the disentanglement capabilities of a state-of-the-art sequence-to-sequence…

机器学习 · 计算机科学 2020-02-18 Vatsal Aggarwal , Marius Cotescu , Nishant Prateek , Jaime Lorenzo-Trueba , Roberto Barra-Chicote

The capability of generating speech with specific type of emotion is desired for many applications of human-computer interaction. Cross-speaker emotion transfer is a common approach to generating emotional speech when speech with emotion…

音频与语音处理 · 电气工程与系统科学 2023-01-05 Guangyan Zhang , Ying Qin , Wenjie Zhang , Jialun Wu , Mei Li , Yutao Gai , Feijun Jiang , Tan Lee

This paper proposes an end-to-end emotional speech synthesis (ESS) method which adopts global style tokens (GSTs) for semi-supervised training. This model is built based on the GST-Tacotron framework. The style tokens are defined to present…

音频与语音处理 · 电气工程与系统科学 2019-06-27 Peng-fei Wu , Zhen-hua Ling , Li-juan Liu , Yuan Jiang , Hong-chuan Wu , Li-rong Dai

Controlling the style and characteristics of speech synthesis is crucial for adapting the output to specific contexts and user requirements. Previous Text-to-speech (TTS) works have focused primarily on the technical aspects of producing…

声音 · 计算机科学 2025-09-04 Jiawei Zhang , Tian-Hao Zhang , Jun Wang , Jiaran Gao , Xinyuan Qian , Xu-Cheng Yin

This paper introduces Easy One-Step Text-to-Speech (E1 TTS), an efficient non-autoregressive zero-shot text-to-speech system based on denoising diffusion pretraining and distribution matching distillation. The training of E1 TTS is…

音频与语音处理 · 电气工程与系统科学 2024-09-17 Zhijun Liu , Shuai Wang , Pengcheng Zhu , Mengxiao Bi , Haizhou Li

State-of-the-art brain-to-text systems have achieved great success in decoding language directly from brain signals using neural networks. However, current approaches are limited to small closed vocabularies which are far from enough for…

人工智能 · 计算机科学 2024-01-09 Zhenhailong Wang , Heng Ji

This paper introduces DiFlow-TTS, a novel zero-shot text-to-speech (TTS) system that employs discrete flow matching for generative speech modeling. We position this work as an entry point that may facilitate further advances in this…

The goal of cross-speaker style transfer in TTS is to transfer a speech style from a source speaker with expressive data to a target speaker with only neutral data. In this context, we propose using a pre-trained singing voice conversion…

音频与语音处理 · 电气工程与系统科学 2024-10-10 Leonardo B. de M. M. Marques , Lucas H. Ueda , Mário U. Neto , Flávio O. Simões , Fernando Runstein , Bianca Dal Bó , Paula D. P. Costa

We propose a novel text-to-speech (TTS) framework centered around a neural transducer. Our approach divides the whole TTS pipeline into semantic-level sequence-to-sequence (seq2seq) modeling and fine-grained acoustic modeling stages,…

音频与语音处理 · 电气工程与系统科学 2024-10-28 Minchan Kim , Myeonghun Jeong , Byoung Jin Choi , Semin Kim , Joun Yeop Lee , Nam Soo Kim

We propose a novel description-based controllable text-to-speech (TTS) method with cross-lingual control capability. To address the lack of audio-description paired data in the target language, we combine a TTS model trained on the target…

音频与语音处理 · 电气工程与系统科学 2024-09-27 Ryuichi Yamamoto , Yuma Shirahata , Masaya Kawamura , Kentaro Tachibana

In expressive speech synthesis, there are high requirements for emotion interpretation. However, it is time-consuming to acquire emotional audio corpus for arbitrary speakers due to their deduction ability. In response to this problem, this…

音频与语音处理 · 电气工程与系统科学 2021-10-12 Pengfei Wu , Junjie Pan , Chenchang Xu , Junhui Zhang , Lin Wu , Xiang Yin , Zejun Ma

Humans can effortlessly modify various prosodic attributes, such as the placement of stress and the intensity of sentiment, to convey a specific emotion while maintaining consistent linguistic content. Motivated by this capability, we…

声音 · 计算机科学 2023-12-29 Leyuan Qu , Wei Wang , Cornelius Weber , Pengcheng Yue , Taihao Li , Stefan Wermter

We present the gradual style adaptor TTS (GSA-TTS) with a novel style encoder that gradually encodes speaking styles from an acoustic reference for zero-shot speech synthesis. GSA first captures the local style of each semantic sound unit.…

计算与语言 · 计算机科学 2025-05-27 Seokgi Lee , Jungjun Kim