中文
相关论文

相关论文: Towards zero-shot Text-based voice editing using a…

200 篇论文

Text-based speech editors expedite the process of editing speech recordings by permitting editing via intuitive cut, copy, and paste operations on a speech transcript. A major drawback of current systems, however, is that edited recordings…

音频与语音处理 · 电气工程与系统科学 2021-02-17 Max Morrison , Lucas Rencker , Zeyu Jin , Nicholas J. Bryan , Juan-Pablo Caceres , Bryan Pardo

Scaling text-to-speech to a large and wild dataset has been proven to be highly effective in achieving timbre and speech style generalization, particularly in zero-shot TTS. However, previous works usually encode speech into latent using…

音频与语音处理 · 电气工程与系统科学 2023-06-07 Ziyue Jiang , Yi Ren , Zhenhui Ye , Jinglin Liu , Chen Zhang , Qian Yang , Shengpeng Ji , Rongjie Huang , Chunfeng Wang , Xiang Yin , Zejun Ma , Zhou Zhao

Motivated by the superior performance of image diffusion models, more and more researchers strive to extend these models to the text-based video editing task. Nevertheless, current video editing tasks mainly suffer from the dilemma between…

计算机视觉与模式识别 · 计算机科学 2023-08-22 Yutao Chen , Xingning Dong , Tian Gan , Chunluan Zhou , Ming Yang , Qingpei Guo

Customizable multilingual zero-shot singing voice synthesis (SVS) has various potential applications in music composition and short video dubbing. However, existing SVS models overly depend on phoneme and note boundary annotations, limiting…

音频与语音处理 · 电气工程与系统科学 2025-08-22 Yu Zhang , Wenxiang Guo , Changhao Pan , Dongyu Yao , Zhiyuan Zhu , Ziyue Jiang , Yuhan Wang , Tao Jin , Zhou Zhao

We introduce StyleFusion-TTS, a prompt and/or audio referenced, style and speaker-controllable, zero-shot text-to-speech (TTS) synthesis system designed to enhance the editability and naturalness of current research literature. We propose a…

音频与语音处理 · 电气工程与系统科学 2024-09-25 Zhiyong Chen , Xinnuo Li , Zhiqi Ai , Shugong Xu

Zero-shot multi-speaker text-to-speech (TTS) systems rely on speaker embeddings to synthesize speech in the voice of an unseen speaker, using only a short reference utterance. While many speaker embeddings have been developed for speaker…

音频与语音处理 · 电气工程与系统科学 2025-09-01 Marie Kunešová , Zdeněk Hanzlíček , Jindřich Matoušek

Zero-shot voice conversion aims to transfer the voice of a source speaker to that of a speaker unseen during training, while preserving the content information. Although various methods have been proposed to reconstruct speaker information…

声音 · 计算机科学 2024-08-22 Anastasia Avdeeva , Aleksei Gusev

Voice conversion (VC) systems are widely used for several applications, from speaker anonymisation to personalised speech synthesis. Supervised approaches learn a mapping between different speakers using parallel data, which is expensive to…

Recent advances in zero-shot text-to-speech (TTS) have enabled accurate imitation of reference speech in terms of both speaking style and speaker timbre. However, achieving disentangled control over these aspects from separate references…

音频与语音处理 · 电气工程与系统科学 2026-05-26 Yoonhyung Lee , Hyunsin Park , Jinhwan Park , Jinkyu Lee

With the popularity of deep neural network, speech synthesis task has achieved significant improvements based on the end-to-end encoder-decoder framework in the recent days. More and more applications relying on speech synthesis technology…

音频与语音处理 · 电气工程与系统科学 2020-10-23 Dongyang Dai , Li Chen , Yuping Wang , Mu Wang , Rui Xia , Xuchen Song , Zhiyong Wu , Yuxuan Wang

Short-utterance speaker verification presents significant challenges due to the limited information in brief speech segments, which can undermine accuracy and reliability. Recently, zero-shot text-to-speech (ZS-TTS) systems have made…

声音 · 计算机科学 2025-06-18 Yiyang Zhao , Shuai Wang , Guangzhi Sun , Zehua Chen , Chao Zhang , Mingxing Xu , Thomas Fang Zheng

Voice style transfer, also called voice conversion, seeks to modify one speaker's voice to generate speech as if it came from another (target) speaker. Previous works have made progress on voice conversion with parallel training data and…

音频与语音处理 · 电气工程与系统科学 2021-03-18 Siyang Yuan , Pengyu Cheng , Ruiyi Zhang , Weituo Hao , Zhe Gan , Lawrence Carin

One-shot voice conversion (VC) aims to convert speech from any source speaker to an arbitrary target speaker with only a few seconds of reference speech from the target speaker. This relies heavily on disentangling the speaker's identity…

音频与语音处理 · 电气工程与系统科学 2023-01-02 Yinghao Aaron Li , Cong Han , Nima Mesgarani

Artificial speech synthesis has made a great leap in terms of naturalness as recent Text-to-Speech (TTS) systems are capable of producing speech with similar quality to human recordings. However, not all speaking styles are easy to model:…

In this work, we propose a zero-shot voice conversion method using speech representations trained with self-supervised learning. First, we develop a multi-task model to decompose a speech utterance into features such as linguistic content,…

声音 · 计算机科学 2023-02-17 Shehzeen Hussain , Paarth Neekhara , Jocelyn Huang , Jason Li , Boris Ginsburg

Zero-shot voice conversion (VC) aims to transfer the timbre from the source speaker to an arbitrary unseen speaker while preserving the original linguistic content. Despite recent advancements in zero-shot VC using language model-based or…

音频与语音处理 · 电气工程与系统科学 2024-12-11 Jixun Yao , Yuguang Yang , Yu Pan , Ziqian Ning , Jiaohao Ye , Hongbin Zhou , Lei Xie

Zero-shot Text-to-Speech (TTS) models can generate speech that captures both the voice timbre and accent of a reference speaker. However, disentangling these attributes remains challenging, as the output often inherits both the accent and…

音频与语音处理 · 电气工程与系统科学 2026-03-09 Mu Yang , John H. L. Hansen

Recently, adapting Vision Language Models (VLMs) to zero-shot visual classification by tuning class embedding with a few prompts (Test-time Prompt Tuning, TPT) or replacing class names with generated visual samples (support-set) has shown…

计算机视觉与模式识别 · 计算机科学 2025-02-11 Rui Yan , Jin Wang , Hongyu Qu , Xiaoyu Du , Dong Zhang , Jinhui Tang , Tieniu Tan

The advancements in zero-shot text-to-speech (TTS) methods, based on large-scale models, have demonstrated high fidelity in reproducing speaker characteristics. However, these models are too large for practical daily use. We propose a…

声音 · 计算机科学 2024-07-02 Kenichi Fujita , Takanori Ashihara , Marc Delcroix , Yusuke Ijima

While neural text-to-speech (TTS) has achieved human-like natural synthetic speech, multilingual TTS systems are limited to resource-rich languages due to the need for paired text and studio-quality audio data. This paper proposes a method…

音频与语音处理 · 电气工程与系统科学 2023-05-30 Takaaki Saeki , Soumi Maiti , Xinjian Li , Shinji Watanabe , Shinnosuke Takamichi , Hiroshi Saruwatari