English
Related papers

Related papers: Voice-preserving Zero-shot Multiple Accent Convers…

200 papers

The cloning of a speaker's voice using an untranscribed reference sample is one of the great advances of modern neural text-to-speech (TTS) methods. Approaches for mimicking the prosody of a transcribed reference audio have also been…

Sound · Computer Science 2022-10-25 Florian Lux , Julia Koch , Ngoc Thang Vu

One-shot voice conversion (VC) with only a single target speaker's speech for reference has become a hot research topic. Existing works generally disentangle timbre, while information about pitch, rhythm and content is still mixed together.…

Audio and Speech Processing · Electrical Eng. & Systems 2022-08-24 SiCheng Yang , Methawee Tantrawenith , Haolin Zhuang , Zhiyong Wu , Aolan Sun , Jianzong Wang , Ning Cheng , Huaizhen Tang , Xintao Zhao , Jie Wang , Helen Meng

The idea of combining multiple languages' recordings to train a single automatic speech recognition (ASR) model brings the promise of the emergence of universal speech representation. Recently, a Transformer encoder-decoder model has been…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-08 Siyuan Feng , Piotr Żelasko , Laureano Moro-Velázquez , Ali Abavisani , Mark Hasegawa-Johnson , Odette Scharenborg , Najim Dehak

The acoustic background plays a crucial role in natural conversation. It provides context and helps listeners understand the environment, but a strong background makes it difficult for listeners to understand spoken words. The appropriate…

Sound · Computer Science 2025-02-12 Leying Zhang , Wangyou Zhang , Zhengyang Chen , Yanmin Qian

Pre-trained multilingual language encoders, such as multilingual BERT and XLM-R, show great potential for zero-shot cross-lingual transfer. However, these multilingual encoders do not precisely align words and phrases across languages.…

Computation and Language · Computer Science 2021-09-13 Kuan-Hao Huang , Wasi Uddin Ahmad , Nanyun Peng , Kai-Wei Chang

Zero-shot voice conversion aims to transform a source speech utterance to match the timbre of a reference speech from an unseen speaker. Traditional approaches struggle with timbre leakage, insufficient timbre representation, and mismatches…

Sound · Computer Science 2024-11-18 Songting Liu

Transfer learning between different language pairs has shown its effectiveness for Neural Machine Translation (NMT) in low-resource scenario. However, existing transfer methods involving a common target language are far from success in the…

Computation and Language · Computer Science 2019-12-04 Baijun Ji , Zhirui Zhang , Xiangyu Duan , Min Zhang , Boxing Chen , Weihua Luo

The human auditory system is able to distinguish the vocal source of thousands of speakers, yet not much is known about what features the auditory system uses to do this. Fourier Transforms are capable of capturing the pitch and harmonic…

Machine Learning · Statistics 2016-10-28 Shariq Mobin , Joan Bruna

We propose a neural network for zero-shot voice conversion (VC) without any parallel or transcribed data. Our approach uses pre-trained models for automatic speech recognition (ASR) and speaker embedding, obtained from a speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-19 Yurii Rebryk , Stanislav Beliaev

Despite their success, large pre-trained multilingual models have not completely alleviated the need for labeled data, which is cumbersome to collect for all target languages. Zero-shot cross-lingual transfer is emerging as a practical…

Computation and Language · Computer Science 2021-07-01 Iulia Turc , Kenton Lee , Jacob Eisenstein , Ming-Wei Chang , Kristina Toutanova

Training deep neural networks for automatic speech recognition (ASR) requires large amounts of transcribed speech. This becomes a bottleneck for training robust models for accented speech which typically contains high variability in…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-11 Nilaksh Das , Sravan Bodapati , Monica Sunkara , Sundararajan Srinivasan , Duen Horng Chau

Artificially generated speech is increasingly embedded in everyday life. Voice cloning in particular enables applications where identity preservation is important, such as completing a recording, dubbing in a new language, or preserving the…

Sound · Computer Science 2026-05-28 Kaitlyn Zhou , Federico Bianchi , Martijn Bartelds , Anna Pot , Yongchan Kwon , James Zou

Voice anonymization has been developed as a technique for preserving privacy by replacing the speaker's voice in a speech signal with that of a pseudo-speaker, thereby obscuring the original voice attributes from machine recognition and…

Sound · Computer Science 2024-11-13 Rui Wang , Liping Chen , Kong AiK Lee , Zhen-Hua Ling

Automatic Speech Recognition (ASR) systems generalize poorly on accented speech. The phonetic and linguistic variability of accents present hard challenges for ASR systems today in both data collection and modeling strategies. The resulting…

In this work we introduce a semi-supervised approach to the voice conversion problem, in which speech from a source speaker is converted into speech of a target speaker. The proposed method makes use of both parallel and non-parallel…

Machine Learning · Statistics 2019-10-02 Cory Stephenson , Gokce Keskin , Anil Thomas , Oguz H. Elibol

Precise control over speech characteristics, such as pitch, duration, and speech rate, remains a significant challenge in the field of voice conversion. The ability to manipulate parameters like pitch and syllable rate is an important…

Sound · Computer Science 2025-07-08 Mathilde Abrassart , Nicolas Obin , Axel Roebel

This work presents self-supervised learning methods for developing monaural speaker-specific (i.e., personalized) speech enhancement models. While generalist models must broadly address many speakers, specialist models can adapt their…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-28 Aswin Sivaraman , Minje Kim

Acoustic word embedding models map variable duration speech segments to fixed dimensional vectors, enabling efficient speech search and discovery. Previous work explored how embeddings can be obtained in zero-resource settings where no…

Computation and Language · Computer Science 2021-06-25 Christiaan Jacobs , Herman Kamper

One-shot style transfer is a challenging task, since training on one utterance makes model extremely easy to over-fit to training data and causes low speaker similarity and lack of expressiveness. In this paper, we build on the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-22 Zhichao Wang , Qicong Xie , Tao Li , Hongqiang Du , Lei Xie , Pengcheng Zhu , Mengxiao Bi

Cross-speaker style transfer aims to extract the speech style of the given reference speech, which can be reproduced in the timbre of arbitrary target speakers. Existing methods on this topic have explored utilizing utterance-level style…

Sound · Computer Science 2022-08-22 Xiang Li , Changhe Song , Xianhao Wei , Zhiyong Wu , Jia Jia , Helen Meng