中文
相关论文

相关论文: Cloning one's voice using very limited data in the…

200 篇论文

In the existing cross-speaker style transfer task, a source speaker with multi-style recordings is necessary to provide the style for a target speaker. However, it is hard for one speaker to express all expected styles. In this paper, a…

音频与语音处理 · 电气工程与系统科学 2021-12-24 Qicong Xie , Tao Li , Xinsheng Wang , Zhichao Wang , Lei Xie , Guoqiao Yu , Guanglu Wan

Recent research in zero-shot speech synthesis has made significant progress in speaker similarity. However, current efforts focus on timbre generalization rather than prosody modeling, which results in limited naturalness and…

声音 · 计算机科学 2024-06-12 Yuepeng Jiang , Tao Li , Fengyu Yang , Lei Xie , Meng Meng , Yujun Wang

Cross-speaker style transfer in speech synthesis aims at transferring a style from source speaker to synthesized speech of a target speaker's timbre. In most previous methods, the synthesized fine-grained prosody features often represent…

声音 · 计算机科学 2023-03-15 Chunyu Qiang , Peng Yang , Hao Che , Ying Zhang , Xiaorui Wang , Zhongyuan Wang

Voice cloning technologies have found applications in a variety of areas ranging from personalized speech interfaces to advertisement, robotics, and so on. Existing voice cloning systems are capable of learning speaker characteristics and…

音频与语音处理 · 电气工程与系统科学 2019-02-20 Hafiz Malik

There are many use cases in singing synthesis where creating voices from small amounts of data is desirable. In text-to-speech there have been several promising results that apply voice cloning techniques to modern deep learning based…

声音 · 计算机科学 2019-02-21 Merlijn Blaauw , Jordi Bonada , Ryunosuke Daido

Speech synthesis has recently seen significant improvements in fidelity, driven by the advent of neural vocoders and neural prosody generators. However, these systems lack intuitive user controls over prosody, making them unable to rectify…

音频与语音处理 · 电气工程与系统科学 2020-08-13 Max Morrison , Zeyu Jin , Justin Salamon , Nicholas J. Bryan , Gautham J. Mysore

We address the problem of human-in-the-loop control for generating prosody in the context of text-to-speech synthesis. Controlling prosody is challenging because existing generative models lack an efficient interface through which users can…

音频与语音处理 · 电气工程与系统科学 2024-04-17 Dan Andrei Iliescu , Devang Savita Ram Mohan , Tian Huey Teh , Zack Hodari

Voice cloning is a highly desired feature for personalized speech interfaces. Neural network based speech synthesis has been shown to generate high quality speech for a large number of speakers. In this paper, we introduce a neural voice…

计算与语言 · 计算机科学 2018-10-15 Sercan O. Arik , Jitong Chen , Kainan Peng , Wei Ping , Yanqi Zhou

Voice cloning is the task of learning to synthesize the voice of an unseen speaker from a few samples. While current voice cloning methods achieve promising results in Text-to-Speech (TTS) synthesis for a new voice, these approaches lack…

声音 · 计算机科学 2021-02-02 Paarth Neekhara , Shehzeen Hussain , Shlomo Dubnov , Farinaz Koushanfar , Julian McAuley

Text does not fully specify the spoken form, so text-to-speech models must be able to learn from speech data that vary in ways not explained by the corresponding text. One way to reduce the amount of unexplained variation in training data…

Artificially generated speech is increasingly embedded in everyday life. Voice cloning in particular enables applications where identity preservation is important, such as completing a recording, dubbing in a new language, or preserving the…

声音 · 计算机科学 2026-05-28 Kaitlyn Zhou , Federico Bianchi , Martijn Bartelds , Anna Pot , Yongchan Kwon , James Zou

The cloning of a speaker's voice using an untranscribed reference sample is one of the great advances of modern neural text-to-speech (TTS) methods. Approaches for mimicking the prosody of a transcribed reference audio have also been…

声音 · 计算机科学 2022-10-25 Florian Lux , Julia Koch , Ngoc Thang Vu

Synthetic-voice cloning technologies have seen significant advances in recent years, giving rise to a range of potential harms. From small- and large-scale financial fraud to disinformation campaigns, the need for reliable methods to…

声音 · 计算机科学 2023-09-28 Sarah Barrington , Romit Barua , Gautham Koorma , Hany Farid

Customizing voice and speaking style in a speech synthesis system with intuitive and fine-grained controls is challenging, given that little data with appropriate labels is available. Furthermore, editing an existing human's voice also…

声音 · 计算机科学 2023-10-27 Florian Lux , Pascal Tilli , Sarina Meyer , Ngoc Thang Vu

Spontaneous style speech synthesis, which aims to generate human-like speech, often encounters challenges due to the scarcity of high-quality data and limitations in model capabilities. Recent language model-based TTS systems can be trained…

声音 · 计算机科学 2024-07-19 Weiqin Li , Peiji Yang , Yicheng Zhong , Yixuan Zhou , Zhisheng Wang , Zhiyong Wu , Xixin Wu , Helen Meng

It is desirable for a text-to-speech system to take into account the environment where synthetic speech is presented, and provide appropriate context-dependent output to the user. In this paper, we present and compare various approaches for…

音频与语音处理 · 电气工程与系统科学 2021-01-15 Qiong Hu , Tobias Bleisch , Petko Petkov , Tuomo Raitio , Erik Marchi , Varun Lakshminarasimhan

In this paper, we present a cross-lingual voice cloning approach. BN features obtained by SI-ASR model are used as a bridge across speakers and language boundaries. The relationships between text and BN features are modeled by the latent…

音频与语音处理 · 电气工程与系统科学 2019-10-31 Xinyong Zhou , Hao Che , Xiaorui Wang , Lei Xie

Recently, sequence-to-sequence (seq-to-seq) models have been successfully applied in text-to-speech (TTS) to synthesize speech for single-language text. To synthesize speech for multiple languages usually requires multi-lingual speech from…

声音 · 计算机科学 2022-11-18 Haitong Zhang , Yue Lin

This study explores voice cloning to generate synthetic speech replicating the unique patterns of individuals with dysarthria. Using the TORGO dataset, we address data scarcity and privacy challenges in speech-language pathology. Our…

声音 · 计算机科学 2025-03-04 Birger Moell , Fredrik Sand Aronsson

Modern neural text-to-speech (TTS) synthesis can generate speech that is indistinguishable from natural speech. However, the prosody of generated utterances often represents the average prosodic style of the database instead of having wide…

音频与语音处理 · 电气工程与系统科学 2020-09-16 Tuomo Raitio , Ramya Rasipuram , Dan Castellani
‹ 上一页 1 2 3 10 下一页 ›