English
Related papers

Related papers: MelGAN-VC: Voice Conversion and Audio Style Transf…

200 papers

Mel-frequency filter bank (MFB) based approaches have the advantage of learning speech compared to raw spectrum since MFB has less feature size. However, speech generator with MFB approaches require additional vocoder that needs a huge…

Audio and Speech Processing · Electrical Eng. & Systems 2020-12-01 June-Woo Kim , Ho-Young Jung , Minho Lee

Zero-shot voice conversion (VC) converts source speech into the voice of any desired speaker using only one utterance of the speaker without requiring additional model updates. Typical methods use a speaker representation from a pre-trained…

Audio and Speech Processing · Electrical Eng. & Systems 2024-05-21 Zhichao Wang , Liumeng Xue , Qiuqiang Kong , Lei Xie , Yuanzhe Chen , Qiao Tian , Yuping Wang

This paper proposes a new task called spatial voice conversion, which aims to convert a target voice while preserving spatial information and non-target signals. Traditional voice conversion methods focus on single-channel waveforms,…

GAN-based neural vocoders, such as Parallel WaveGAN and MelGAN have attracted great interest due to their lightweight and parallel structures, enabling them to generate high fidelity waveform in a real-time manner. In this paper, inspired…

Sound · Computer Science 2021-03-30 Congyi Wang , Yu Chen , Bin Wang , Yi Shi

Generative audio technologies now enable highly realistic voice cloning and real-time voice conversion, increasing the risk of impersonation, fraud, and misinformation in communication channels such as phone and video calls. This study…

Sound · Computer Science 2026-01-09 Prajwal Chinchmalatpure , Suyash Chinchmalatpure , Siddharth Chavan

Modern speaker recognition system relies on abundant and balanced datasets for classification training. However, diverse defective datasets, such as partially-labelled, small-scale, and imbalanced datasets, are common in real-world…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-03 Ruijie Tao , Zhan Shi , Yidi Jiang , Tianchi Liu , Haizhou Li

For training the sequence-to-sequence voice conversion model, we need to handle an issue of insufficient data about the number of speech pairs which consist of the same utterance. This study experimentally investigated the effects of…

Machine Learning · Computer Science 2020-06-16 Yeongtae Hwang , Hyemin Cho , Hongsun Yang , Dong-Ok Won , Insoo Oh , Seong-Whan Lee

Zero-shot voice conversion aims to transfer the voice of a source speaker to that of a speaker unseen during training, while preserving the content information. Although various methods have been proposed to reconstruct speaker information…

Sound · Computer Science 2024-08-22 Anastasia Avdeeva , Aleksei Gusev

Zero-Shot Voice Conversion (VC) aims to transform the source speaker's timbre into an arbitrary unseen one while retaining speech content. Most prior work focuses on preserving the source's prosody, while fine-grained timbre information may…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-03 Jialong Zuo , Shengpeng Ji , Minghui Fang , Mingze Li , Ziyue Jiang , Xize Cheng , Xiaoda Yang , Chen Feiyang , Xinyu Duan , Zhou Zhao

In addition to conveying the linguistic content from source speech to converted speech, maintaining the speaking style of source speech also plays an important role in the voice conversion (VC) task, which is essential in many scenarios…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-06 Zhichao Wang , Xinsheng Wang , Qicong Xie , Tao Li , Lei Xie , Qiao Tian , Yuping Wang

Cross-lingual voice conversion (VC) is an important and challenging problem due to significant mismatches of the phonetic set and the speech prosody of different languages. In this paper, we build upon the neural text-to-speech (TTS) model,…

Sound · Computer Science 2021-02-04 Shengkui Zhao , Hao Wang , Trung Hieu Nguyen , Bin Ma

Privacy-preserving voice conversion aims to remove only the attributes of speech audio that convey identity information, keeping other speech characteristics intact. This paper presents a mechanism for privacy-preserving voice conversion…

Sound · Computer Science 2024-09-24 Jacob J Webber , Oliver Watts , Gustav Eje Henter , Jennifer Williams , Simon King

This paper proposes an interesting voice and accent joint conversion approach, which can convert an arbitrary source speaker's voice to a target speaker with non-native accent. This problem is challenging as each target speaker only has…

Sound · Computer Science 2020-11-18 Zhichao Wang , Wenshuo Ge , Xiong Wang , Shan Yang , Wendong Gan , Haitao Chen , Hai Li , Lei Xie , Xiulin Li

In this paper, we describe our speech generation system for the first Audio Deep Synthesis Detection Challenge (ADD 2022). Firstly, we build an any-to-many voice conversion (VC) system to convert source speech with arbitrary language…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-21 Cheng Wen , Tingwei Guo , Xingjun Tan , Rui Yan , Shuran Zhou , Chuandong Xie , Wei Zou , Xiangang Li

We introduce HybridVC, a voice conversion (VC) framework built upon a pre-trained conditional variational autoencoder (CVAE) that combines the strengths of a latent model with contrastive learning. HybridVC supports text and audio prompts,…

Sound · Computer Science 2024-09-26 Xinlei Niu , Jing Zhang , Charles Patrick Martin

Style transfer is a technique for combining two images based on the activations and feature statistics in a deep learning neural network architecture. This paper studies the analogous task in the audio domain and takes a critical look at…

Sound · Computer Science 2020-08-10 M. Huzaifah , L. Wyse

Speech enhancement is an essential task of improving speech quality in noise scenario. Several state-of-the-art approaches have introduced visual information for speech enhancement,since the visual aspect of speech is essentially unaffected…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-21 Xinmeng Xu , Yang Wang , Dongxiang Xu , Yiyuan Peng , Cong Zhang , Jie Jia , Binbin Chen

One-shot voice conversion (VC) with only a single target speaker's speech for reference has become a hot research topic. Existing works generally disentangle timbre, while information about pitch, rhythm and content is still mixed together.…

Audio and Speech Processing · Electrical Eng. & Systems 2022-08-24 SiCheng Yang , Methawee Tantrawenith , Haolin Zhuang , Zhiyong Wu , Aolan Sun , Jianzong Wang , Ning Cheng , Huaizhen Tang , Xintao Zhao , Jie Wang , Helen Meng

This paper presents an adversarial learning method for recognition-synthesis based non-parallel voice conversion. A recognizer is used to transform acoustic features into linguistic representations while a synthesizer recovers output…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-07 Jing-Xuan Zhang , Zhen-Hua Ling , Li-Rong Dai

Advanced Generative Adversarial Networks (GANs) are remarkable in generating intelligible audio from a random latent vector. In this paper, we examine the task of recovering the latent vector of both synthesized and real audio. Previous…

Sound · Computer Science 2020-10-19 Andrew Keyes , Nicky Bayat , Vahid Reza Khazaie , Yalda Mohsenzadeh
‹ Prev 1 8 9 10 Next ›