English
Related papers

Related papers: Fine-grained robust prosody transfer for single-sp…

200 papers

We present an extension to the Tacotron speech synthesis architecture that learns a latent embedding space of prosody, derived from a reference acoustic representation containing the desired prosody. We show that conditioning Tacotron on…

Computation and Language · Computer Science 2018-03-28 RJ Skerry-Ryan , Eric Battenberg , Ying Xiao , Yuxuan Wang , Daisy Stanton , Joel Shor , Ron J. Weiss , Rob Clark , Rif A. Saurous

We propose prosody embeddings for emotional and expressive speech synthesis networks. The proposed methods introduce temporal structures in the embedding networks, thus enabling fine-grained control of the speaking style of the synthesized…

Computation and Language · Computer Science 2019-02-19 Younggun Lee , Taesu Kim

Text-to-speech systems recently achieved almost indistinguishable quality from human speech. However, the prosody of those systems is generally flatter than natural speech, producing samples with low expressiveness. Disentanglement of…

Prosody Transfer (PT) is a technique that aims to use the prosody from a source audio as a reference while synthesising speech. Fine-grained PT aims at capturing prosodic aspects like rhythm, emphasis, melody, duration, and loudness, from a…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-11 Sri Karlapati , Alexis Moinet , Arnaud Joly , Viacheslav Klimkov , Daniel Sáez-Trigueros , Thomas Drugman

Some recent models for Text-to-Speech synthesis aim to transfer the prosody of a reference utterance to the generated target synthetic speech. This is done by using a learned embedding of the reference utterance, which is used to condition…

Computation and Language · Computer Science 2023-03-09 Atli Thor Sigurgeirsson , Simon King

This paper presents a simple yet effective method to achieve prosody transfer from a reference speech signal to synthesized speech. The main idea is to incorporate well-known acoustic correlates of prosody such as pitch and loudness…

Sound · Computer Science 2020-05-19 Siddharth Gururani , Kilol Gupta , Dhaval Shah , Zahra Shakeri , Jervis Pinto

Prosody transfer is well-studied in the context of expressive speech synthesis. Cross-lingual prosody transfer, however, is challenging and has been under-explored to date. In this paper, we present a novel solution to learn prosody…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-21 Jakub Swiatkowski , Duo Wang , Mikolaj Babianski , Patrick Lumban Tobing , Ravichander Vipperla , Vincent Pollet

Some recent studies have demonstrated the feasibility of single-stage neural text-to-speech, which does not need to generate mel-spectrograms but generates the raw waveforms directly from the text. Single-stage text-to-speech often faces…

Sound · Computer Science 2022-07-14 Zhengxi Liu , Qiao Tian , Chenxu Hu , Xudong Liu , Menglin Wu , Yuping Wang , Hang Zhao , Yuxuan Wang

Speech-to-speech translation systems today do not adequately support use for dialog purposes. In particular, nuances of speaker intent and stance can be lost due to improper prosody transfer. We present an exploration of what needs to be…

Computation and Language · Computer Science 2023-07-11 Jonathan E. Avila , Nigel G. Ward

A text-to-speech (TTS) model typically factorizes speech attributes such as content, speaker and prosody into disentangled representations.Recent works aim to additionally model the acoustic conditions explicitly, in order to disentangle…

Zero-shot speaker adaptation aims to clone an unseen speaker's voice without any adaptation time and parameters. Previous researches usually use a speaker encoder to extract a global fixed speaker embedding from reference speech, and…

Sound · Computer Science 2022-11-14 Yixuan Zhou , Changhe Song , Xiang Li , Luwen Zhang , Zhiyong Wu , Yanyao Bian , Dan Su , Helen Meng

The cloning of a speaker's voice using an untranscribed reference sample is one of the great advances of modern neural text-to-speech (TTS) methods. Approaches for mimicking the prosody of a transcribed reference audio have also been…

Sound · Computer Science 2022-10-25 Florian Lux , Julia Koch , Ngoc Thang Vu

Cross-speaker style transfer is crucial to the applications of multi-style and expressive speech synthesis at scale. It does not require the target speakers to be experts in expressing all styles and to collect corresponding recordings for…

Sound · Computer Science 2021-07-28 Shifeng Pan , Lei He

Modern neural TTS systems are capable of generating natural and expressive speech when provided with sufficient amounts of training data. Such systems can be equipped with prosody-control functionality, allowing for more direct shaping of…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-21 Slava Shechtman , Raul Fernandez

This paper presents a novel design of neural network system for fine-grained style modeling, transfer and prediction in expressive text-to-speech (TTS) synthesis. Fine-grained modeling is realized by extracting style embeddings from the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-11 Daxin Tan , Tan Lee

Text-to-speech is now able to achieve near-human naturalness and research focus has shifted to increasing expressivity. One popular method is to transfer the prosody from a reference speech sample. There have been considerable advances in…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-22 Alexandra Torresquintero , Tian Huey Teh , Christopher G. R. Wallis , Marlene Staib , Devang S Ram Mohan , Vivian Hu , Lorenzo Foglianti , Jiameng Gao , Simon King

Textless speech-to-speech translation systems are rapidly advancing, thanks to the integration of self-supervised learning techniques. However, existing state-of-the-art systems fall short when it comes to capturing and transferring…

Sound · Computer Science 2023-10-12 Jarod Duret , Benjamin O'Brien , Yannick Estève , Titouan Parcollet

In the existing cross-speaker style transfer task, a source speaker with multi-style recordings is necessary to provide the style for a target speaker. However, it is hard for one speaker to express all expected styles. In this paper, a…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-24 Qicong Xie , Tao Li , Xinsheng Wang , Zhichao Wang , Lei Xie , Guoqiao Yu , Guanglu Wan

Current voice conversion (VC) methods can successfully convert timbre of the audio. As modeling source audio's prosody effectively is a challenging task, there are still limitations of transferring source style to the converted speech. This…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-29 Zhichao Wang , Xinyong Zhou , Fengyu Yang , Tao Li , Hongqiang Du , Lei Xie , Wendong Gan , Haitao Chen , Hai Li

Modern sequence to sequence neural TTS systems provide close to natural speech quality. Such systems usually comprise a network converting linguistic/phonetic features sequence to an acoustic features sequence, cascaded with a neural…

Audio and Speech Processing · Electrical Eng. & Systems 2019-09-26 Slava Shechtman , Alex Sorin
‹ Prev 1 2 3 10 Next ›