中文
相关论文

相关论文: Multi-speaker Multi-style Text-to-speech Synthesis…

200 篇论文

The recent progress in non-autoregressive text-to-speech (NAR-TTS) has made fast and high-quality speech synthesis possible. However, current NAR-TTS models usually use phoneme sequence as input and thus cannot understand the…

声音 · 计算机科学 2022-04-26 Zhenhui Ye , Zhou Zhao , Yi Ren , Fei Wu

Though significant progress has been made for speaker-dependent Video-to-Speech (VTS) synthesis, little attention is devoted to multi-speaker VTS that can map silent video to speech, while allowing flexible control of speaker identity, all…

音频与语音处理 · 电气工程与系统科学 2022-02-21 Disong Wang , Shan Yang , Dan Su , Xunying Liu , Dong Yu , Helen Meng

Expressive speech synthesis is crucial for many human-computer interaction scenarios, such as audiobooks, podcasts, and voice assistants. Previous works focus on predicting the style embeddings at one single scale from the information…

声音 · 计算机科学 2023-08-01 Shun Lei , Yixuan Zhou , Liyang Chen , Zhiyong Wu , Xixin Wu , Shiyin Kang , Helen Meng

Currently, a common approach in many speech processing tasks is to leverage large scale pre-trained models by fine-tuning them on in-domain data for a particular application. Yet obtaining even a small amount of such data can be…

音频与语音处理 · 电气工程与系统科学 2024-08-20 Samuele Cornell , Jordan Darefsky , Zhiyao Duan , Shinji Watanabe

Expressive text-to-speech (TTS) can synthesize a new speaking style by imiating prosody and timbre from a reference audio, which faces the following challenges: (1) The highly dynamic prosody information in the reference audio is difficult…

声音 · 计算机科学 2022-11-07 Dongchao Yang , Songxiang Liu , Jianwei Yu , Helin Wang , Chao Weng , Yuexian Zou

This paper introduces Taco-VC, a novel architecture for voice conversion based on Tacotron synthesizer, which is a sequence-to-sequence with attention model. The training of multi-speaker voice conversion systems requires a large number of…

声音 · 计算机科学 2020-06-22 Roee Levy Leshem , Raja Giryes

Text-to-speech (TTS) systems that scale up the amount of training data have achieved significant improvements in zero-shot speech synthesis. However, these systems have certain limitations: they require a large amount of training data,…

音频与语音处理 · 电气工程与系统科学 2024-10-07 Taejun Bak , Youngsik Eom , SeungJae Choi , Young-Sun Joo

By representing speaker characteristic as a single fixed-length vector extracted solely from speech, we can train a neural multi-speaker speech synthesis model by conditioning the model on those vectors. This model can also be adapted to…

音频与语音处理 · 电气工程与系统科学 2019-10-09 Hieu-Thi Luong , Junichi Yamagishi

The style of the speech varies from person to person and every person exhibits his or her own style of speaking that is determined by the language, geography, culture and other factors. Style is best captured by prosody of a signal. High…

音频与语音处理 · 电气工程与系统科学 2020-12-15 Neeraj Kumar , Srishti Goel , Ankur Narang , Brejesh Lall

Conventional speech-to-text translation (ST) systems are trained on single-speaker utterances, and they may not generalize to real-life scenarios where the audio contains conversations by multiple speakers. In this paper, we tackle…

Stylistic text generation plays a vital role in enhancing communication by reflecting the nuances of individual expression. This paper presents a novel approach for generating text in a specific speaker's style across different languages.…

计算与语言 · 计算机科学 2025-01-23 Karishma Thakrar , Katrina Lawrence , Kyle Howard

For articulatory-to-acoustic mapping, typically only limited parallel training data is available, making it impossible to apply fully end-to-end solutions like Tacotron2. In this paper, we experimented with transfer learning and adaptation…

音频与语音处理 · 电气工程与系统科学 2021-07-27 Csaba Zainkó , László Tóth , Amin Honarmandi Shandiz , Gábor Gosztolya , Alexandra Markó , Géza Németh , Tamás Gábor Csapó

In addition to conveying the linguistic content from source speech to converted speech, maintaining the speaking style of source speech also plays an important role in the voice conversion (VC) task, which is essential in many scenarios…

音频与语音处理 · 电气工程与系统科学 2023-09-06 Zhichao Wang , Xinsheng Wang , Qicong Xie , Tao Li , Lei Xie , Qiao Tian , Yuping Wang

Cross-lingual synthesis can be defined as the task of letting a speaker generate fluent synthetic speech in another language. This is a challenging task, and resulting speech can suffer from reduced naturalness, accented speech, and/or loss…

声音 · 计算机科学 2022-04-04 Marcel de Korte , Jaebok Kim , Aki Kunikoshi , Adaeze Adigwe , Esther Klabbers

Multi-speaker spoken datasets enable the creation of text-to-speech synthesis (TTS) systems which can output several voice identities. The multi-speaker (MSPK) scenario also enables the use of fewer training samples per speaker. However, in…

音频与语音处理 · 电气工程与系统科学 2021-06-04 Beata Lorincz , Adriana Stan , Mircea Giurgiu

Building a high-quality singing corpus for a person who is not good at singing is non-trivial, thus making it challenging to create a singing voice synthesizer for this person. Learn2Sing is dedicated to synthesizing the singing voice of a…

声音 · 计算机科学 2022-05-27 Heyang Xue , Xinsheng Wang , Yongmao Zhang , Lei Xie , Pengcheng Zhu , Mengxiao Bi

The front-end is a critical component of English text-to-speech (TTS) systems, responsible for extracting linguistic features that are essential for a text-to-speech model to synthesize speech, such as prosodies and phonemes. The English…

计算与语言 · 计算机科学 2024-03-27 Zelin Ying , Chen Li , Yu Dong , Qiuqiang Kong , Qiao Tian , Yuanyuan Huo , Yuxuan Wang

In a typical voice conversion system, prior works utilize various acoustic features (e.g., the pitch, voiced/unvoiced flag, aperiodicity) of the source speech to control the prosody of generated waveform. However, the prosody is related…

声音 · 计算机科学 2020-06-01 Zheng Lian , Zhengqi Wen

We explore pretraining strategies including choice of base corpus with the aim of choosing the best strategy for zero-shot multi-speaker end-to-end synthesis. We also examine choice of neural vocoder for waveform synthesis, as well as…

声音 · 计算机科学 2020-11-11 Erica Cooper , Xin Wang , Yi Zhao , Yusuke Yasuda , Junichi Yamagishi

In the development of neural text-to-speech systems, model pre-training with a large amount of non-target speakers' data is a common approach. However, in terms of ultimately achieved system performance for target speaker(s), the actual…

音频与语音处理 · 电气工程与系统科学 2021-10-11 Guangyan Zhang , Yichong Leng , Daxin Tan , Ying Qin , Kaitao Song , Xu Tan , Sheng Zhao , Tan Lee