English
Related papers

Related papers: Zero-shot text-to-speech synthesis conditioned usi…

200 papers

Few-shot speaker adaptation is a specific Text-to-Speech (TTS) system that aims to reproduce a novel speaker's voice with a few training data. While numerous attempts have been made to the few-shot speaker adaptation system, there is still…

Audio and Speech Processing · Electrical Eng. & Systems 2021-08-17 Ji-Hoon Kim , Sang-Hoon Lee , Ji-Hyun Lee , Hong-Gyu Jung , Seong-Whan Lee

Text to speech (TTS) and automatic speech recognition (ASR) are two dual tasks in speech processing and both achieve impressive performance thanks to the recent advance in deep learning and large amount of aligned speech and text data.…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-28 Yi Ren , Xu Tan , Tao Qin , Sheng Zhao , Zhou Zhao , Tie-Yan Liu

The rapid development of neural text-to-speech (TTS) systems enabled its usage in other areas of natural language processing such as automatic speech recognition (ASR) or spoken language translation (SLT). Due to the large number of…

Computation and Language · Computer Science 2024-08-01 Nick Rossenbach , Ralf Schlüter , Sakriani Sakti

Scaling text-to-speech (TTS) to large-scale, multi-speaker, and in-the-wild datasets is important to capture the diversity in human speech such as speaker identities, prosodies, and styles (e.g., singing). Current large TTS systems usually…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-31 Kai Shen , Zeqian Ju , Xu Tan , Yanqing Liu , Yichong Leng , Lei He , Tao Qin , Sheng Zhao , Jiang Bian

We introduce a language modeling approach for text to speech synthesis (TTS). Specifically, we train a neural codec language model (called Vall-E) using discrete codes derived from an off-the-shelf neural audio codec model, and regard TTS…

Computation and Language · Computer Science 2023-01-06 Chengyi Wang , Sanyuan Chen , Yu Wu , Ziqiang Zhang , Long Zhou , Shujie Liu , Zhuo Chen , Yanqing Liu , Huaming Wang , Jinyu Li , Lei He , Sheng Zhao , Furu Wei

In this study, we address the importance of modeling behavior style in virtual agents for personalized human-agent interaction. We propose a machine learning approach to synthesize gestures, driven by prosodic features and text, in the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-23 Mireille Fares , Catherine Pelachaud , Nicolas Obin

Neural Text-to-speech (TTS) synthesis is a powerful technology that can generate speech using neural networks. One of the most remarkable features of TTS synthesis is its capability to produce speech in the voice of different speakers. This…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-19 Vinotha R , Hepsiba D , L. D. Vijay Anand , Deepak John Reji

We propose a method for the task of text-conditioned speech insertion, i.e. inserting a speech sample in an input speech sample, conditioned on the corresponding complete text transcript. An example use case of the task would be to update…

Sound · Computer Science 2025-08-26 Neeraj Matiyali , Siddharth Srivastava , Gaurav Sharma

We present SALAD, a zero-shot TTS autoregressive model operating over continuous speech representations. SALAD utilizes a per-token diffusion process to refine and predict continuous representations for the next time step. We compare our…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-25 Arnon Turetzky , Avihu Dekel , Nimrod Shabtay , Slava Shechtman , David Haws , Hagai Aronowitz , Ron Hoory , Yossi Adi

In this paper, we propose a novel prosody disentangle method for prosodic Text-to-Speech (TTS) model, which introduces the vector quantization (VQ) method to the auxiliary prosody encoder to obtain the decomposed prosody representations in…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-08 Yutian Wang , Yuankun Xie , Kun Zhao , Hui Wang , Qin Zhang

In this paper, we introduce a zero-shot Voice Transfer (VT) module that can be seamlessly integrated into a multi-lingual Text-to-speech (TTS) system to transfer an individual's voice across languages. Our proposed VT module comprises a…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-24 Fadi Biadsy , Youzheng Chen , Isaac Elias , Kyle Kastner , Gary Wang , Andrew Rosenberg , Bhuvana Ramabhadran

Self-supervised learning (SSL) methods which learn representations of data without explicit supervision have gained popularity in speech-processing tasks, particularly for single-talker applications. However, these models often have…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-02 Zili Huang , Desh Raj , Paola García , Sanjeev Khudanpur

Existing autoregressive large-scale text-to-speech (TTS) models have advantages in speech naturalness, but their token-by-token generation mechanism makes it difficult to precisely control the duration of synthesized speech. This becomes a…

Computation and Language · Computer Science 2025-09-04 Siyi Zhou , Yiquan Zhou , Yi He , Xun Zhou , Jinchao Wang , Wei Deng , Jingchen Shu

Spontaneous speaking style exhibits notable differences from other speaking styles due to various spontaneous phenomena (e.g., filled pauses, prolongation) and substantial prosody variation (e.g., diverse pitch and duration variation,…

Sound · Computer Science 2024-01-09 Hanzhao Li , Xinfa Zhu , Liumeng Xue , Yang Song , Yunlin Chen , Lei Xie

Recent advancements in Text-to-Speech (TTS) systems have enabled the generation of natural and expressive speech from textual input. Accented TTS aims to enhance user experience by making the synthesized speech more relatable to minority…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-18 Jan Melechovsky , Ambuj Mehrish , Berrak Sisman , Dorien Herremans

While recent text-to-speech (TTS) systems have made remarkable strides toward human-level quality, the performance of cross-lingual TTS lags behind that of intra-lingual TTS. This gap is mainly rooted from the speaker-language entanglement…

Sound · Computer Science 2023-06-13 Ji-Hoon Kim , Hong-Sun Yang , Yoon-Cheol Ju , Il-Hwan Kim , Byeong-Yeol Kim

Contextual automatic speech recognition (ASR) systems allow for recognizing out-of-vocabulary (OOV) words, such as named entities or rare words. However, it remains challenging due to limited training data and ambiguous or inconsistent…

Computation and Language · Computer Science 2025-09-03 Changsong Liu , Yizhou Peng , Eng Siong Chng

Vocoders received renewed attention as main components in statistical parametric text-to-speech (TTS) synthesis and speech transformation systems. Even though there are vocoding techniques give almost accepted synthesized speech, their high…

Sound · Computer Science 2021-06-22 Mohammed Salah Al-Radhi , Tamás Gábor Csapó , Géza Németh

The cloning of a speaker's voice using an untranscribed reference sample is one of the great advances of modern neural text-to-speech (TTS) methods. Approaches for mimicking the prosody of a transcribed reference audio have also been…

Sound · Computer Science 2022-10-25 Florian Lux , Julia Koch , Ngoc Thang Vu

While controllable Text-to-Speech (TTS) has achieved notable progress, most existing methods remain limited to inter-utterance-level control, making fine-grained intra-utterance expression challenging due to their reliance on non-public…

Sound · Computer Science 2026-05-19 Qifan Liang , Yuansen Liu , Ruixin Wei , Nan Lu , Junchuan Zhao , Ye Wang