English
Related papers

Related papers: Expressive TTS Training with Frame and Style Recon…

200 papers

In this paper, we propose a neural articulation-to-speech (ATS) framework that synthesizes high-quality speech from articulatory signal in a multi-speaker situation. Most conventional ATS approaches only focus on modeling contextual…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-22 Miseul Kim , Zhenyu Piao , Jihyun Lee , Hong-Goo Kang

Although neural end-to-end text-to-speech models can synthesize highly natural speech, there is still room for improvements to its efficiency and naturalness. This paper proposes a non-autoregressive neural text-to-speech model augmented…

Sound · Computer Science 2020-10-23 Isaac Elias , Heiga Zen , Jonathan Shen , Yu Zhang , Ye Jia , Ron Weiss , Yonghui Wu

Neural end-to-end TTS can generate very high-quality synthesized speech, and even close to human recording within similar domain text. However, it performs unsatisfactory when scaling it to challenging test sets. One concern is that the…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-23 Yibin Zheng , Xi Wang , Lei He , Shifeng Pan , Frank K. Soong , Zhengqi Wen , Jianhua Tao

Controlling speaking style in text-to-speech (TTS) systems has become a growing focus in both academia and industry. While many existing approaches rely on reference audio to guide style generation, such methods are often impractical due to…

Sound · Computer Science 2025-10-22 Haowei Lou , Hye-Young Paik , Wen Hu , Lina Yao

Automatic detection of prominence at the word and syllable-levels is critical for building computer-assisted language learning systems. It has been shown that prosody embeddings learned by the current state-of-the-art (SOTA) text-to-speech…

Computation and Language · Computer Science 2024-12-12 Anindita Mondal , Rangavajjala Sankara Bharadwaj , Jhansi Mallela , Anil Kumar Vuppala , Chiranjeevi Yarra

In this paper, we investigate the semi-supervised joint training of text to speech (TTS) and automatic speech recognition (ASR), where a small amount of paired data and a large amount of unpaired text data are available. Conventional…

Sound · Computer Science 2022-07-12 Naoki Makishima , Satoshi Suzuki , Atsushi Ando , Ryo Masumura

This letter presents an incremental text-to-speech (TTS) method that performs synthesis in small linguistic units while maintaining the naturalness of output speech. Incremental TTS is generally subject to a trade-off between latency and…

Sound · Computer Science 2021-05-26 Takaaki Saeki , Shinnosuke Takamichi , Hiroshi Saruwatari

Dysarthric speech exhibits high variability and limited labeled data, posing major challenges for both automatic speech recognition (ASR) and assistive speech technologies. Existing approaches rely on synthetic data augmentation or speech…

Expressive speech synthesis, like audiobook synthesis, is still challenging for style representation learning and prediction. Deriving from reference audio or predicting style tags from text requires a huge amount of labeled data, which is…

Sound · Computer Science 2022-06-28 Yihan Wu , Xi Wang , Shaofei Zhang , Lei He , Ruihua Song , Jian-Yun Nie

While prompt-based text-to-speech (TTS) models enable natural language-driven speaking style control, they often provide limited fine-grained control and apply a single global style across an utterance. This restricts practical use cases…

Computation and Language · Computer Science 2026-05-28 Jaehoon Kang , Yejin Lee , Yoonji Park , Kyuhong Shim

The goal of expressive Text-to-speech (TTS) is to synthesize natural speech with desired content, prosody, emotion, or timbre, in high expressiveness. Most of previous studies attempt to generate speech from given labels of styles and…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-29 Jianhong Tu , Zeyu Cui , Xiaohuan Zhou , Siqi Zheng , Kai Hu , Ju Fan , Chang Zhou

Text-to-speech (TTS) acoustic models map linguistic features into an acoustic representation out of which an audible waveform is generated. The latest and most natural TTS systems build a direct mapping between linguistic and waveform…

Sound · Computer Science 2019-09-24 David Álvarez , Santiago Pascual , Antonio Bonafonte

In this work we propose a novel token-based training strategy that improves Transformer-Transducer (T-T) based speaker change detection (SCD) performance. The conventional T-T based SCD model loss optimizes all output tokens equally. Due to…

Audio and Speech Processing · Electrical Eng. & Systems 2022-12-06 Guanlong Zhao , Quan Wang , Han Lu , Yiling Huang , Ignacio Lopez Moreno

Text-to-speech (TTS) methods have shown promising results in voice cloning, but they require a large number of labeled text-speech pairs. Minimally-supervised speech synthesis decouples TTS by combining two types of discrete speech…

Sound · Computer Science 2023-12-19 Chunyu Qiang , Hao Li , Yixin Tian , Yi Zhao , Ying Zhang , Longbiao Wang , Jianwu Dang

Incorporating cross-speaker style transfer in text-to-speech (TTS) models is challenging due to the need to disentangle speaker and style information in audio. In low-resource expressive data scenarios, voice conversion (VC) can generate…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-27 Lucas H. Ueda , Leonardo B. de M. M. Marques , Flávio O. Simões , Mário U. Neto , Fernando Runstein , Bianca Dal Bó , Paula D. P. Costa

Speech pre-training has primarily demonstrated efficacy on classification tasks, while its capability of generating novel speech, similar to how GPT-2 can generate coherent paragraphs, has barely been explored. Generative Spoken Language…

Training a multi-speaker Text-to-Speech (TTS) model from scratch is computationally expensive and adding new speakers to the dataset requires the model to be re-trained. The naive solution of sequential fine-tuning of a model for new…

Computation and Language · Computer Science 2022-04-01 Hamed Hemati , Damian Borth

This paper proposes an effective emotional text-to-speech (TTS) system with a pre-trained language model (LM)-based emotion prediction method. Unlike conventional systems that require auxiliary inputs such as manually defined emotion…

Current voice conversion (VC) methods can successfully convert timbre of the audio. As modeling source audio's prosody effectively is a challenging task, there are still limitations of transferring source style to the converted speech. This…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-29 Zhichao Wang , Xinyong Zhou , Fengyu Yang , Tao Li , Hongqiang Du , Lei Xie , Wendong Gan , Haitao Chen , Hai Li

Emotion embedding space learned from references is a straightforward approach for emotion transfer in encoder-decoder structured emotional text to speech (TTS) systems. However, the transferred emotion in the synthetic speech is not…

Sound · Computer Science 2020-11-18 Tao Li , Shan Yang , Liumeng Xue , Lei Xie