English
Related papers

Related papers: Discrete Unit based Masking for Improving Disentan…

200 papers

This paper proposes an interesting voice and accent joint conversion approach, which can convert an arbitrary source speaker's voice to a target speaker with non-native accent. This problem is challenging as each target speaker only has…

Sound · Computer Science 2020-11-18 Zhichao Wang , Wenshuo Ge , Xiong Wang , Shan Yang , Wendong Gan , Haitao Chen , Hai Li , Lei Xie , Xiulin Li

Expressive voice conversion (VC) conducts speaker identity conversion for emotional speakers by jointly converting speaker identity and emotional style. Emotional style modeling for arbitrary speakers in expressive VC has not been…

Audio and Speech Processing · Electrical Eng. & Systems 2024-05-06 Zongyang Du , Junchen Lu , Kun Zhou , Lakshmish Kaushik , Berrak Sisman

The task of video-to-speech aims to translate silent video of lip movement to its corresponding audio signal. Previous approaches to this task are generally limited to the case of a single speaker, but a method that accounts for multiple…

Audio and Speech Processing · Electrical Eng. & Systems 2021-05-21 Dan Oneata , Adriana Stan , Horia Cucu

Recently end-to-end neural audio/speech coding has shown its great potential to outperform traditional signal analysis based audio codecs. This is mostly achieved by following the VQ-VAE paradigm where blind features are learned,…

Sound · Computer Science 2023-02-28 Xue Jiang , Xiulian Peng , Yuan Zhang , Yan Lu

We introduce DISSC, a novel, lightweight method that converts the rhythm, pitch contour and timbre of a recording to a target speaker in a textless manner. Unlike DISSC, most voice conversion (VC) methods focus primarily on timbre, and…

Sound · Computer Science 2023-10-20 Gallil Maimon , Yossi Adi

Zero-shot voice conversion is a technique that alters the speaker identity of an input speech to match a target speaker using only a single reference utterance, without requiring additional training. Recent approaches extensively utilize…

Sound · Computer Science 2025-09-11 Youngjun Sim , Jinsung Yoon , Wooyeol Jeong , Young-Joo Suh

The objective of deep learning methods based on encoder-decoder architectures for music source separation is to approximate either ideal time-frequency masks or spectral representations of the target music source(s). The spectral…

We propose a flexible framework that deals with both singer conversion and singers vocal technique conversion. The proposed model is trained on non-parallel corpora, accommodates many-to-many conversion, and leverages recent advances of…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-26 Yin-Jyun Luo , Chin-Chen Hsu , Kat Agres , Dorien Herremans

One-shot voice conversion (VC), which performs conversion across arbitrary speakers with only a single target-speaker utterance for reference, can be effectively achieved by speech representation disentanglement. Existing work generally…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-22 Disong Wang , Liqun Deng , Yu Ting Yeung , Xiao Chen , Xunying Liu , Helen Meng

We propose SelfVC, a training strategy to iteratively improve a voice conversion model with self-synthesized examples. Previous efforts on voice conversion focus on factorizing speech into explicitly disentangled representations that…

The performance of speech enhancement algorithms in a multi-speaker scenario depends on correctly identifying the target speaker to be enhanced. Auditory attention decoding (AAD) methods allow to identify the target speaker which the…

Sound · Computer Science 2020-05-12 Ali Aroudi , Marc Delcroix , Tomohiro Nakatani , Keisuke Kinoshita , Shoko Araki , Simon Doclo

Voice conversion is the task to transform voice characteristics of source speech while preserving content information. Nowadays, self-supervised representation learning models are increasingly utilized in content extraction. However, in…

Sound · Computer Science 2024-05-02 Yimin Deng , Jianzong Wang , Xulong Zhang , Ning Cheng , Jing Xiao

Factorizing speech as disentangled speech representations is vital to achieve highly controllable style transfer in voice conversion (VC). Conventional speech representation learning methods in VC only factorize speech as speaker and…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-06 Jie Wang , Jingbei Li , Xintao Zhao , Zhiyong Wu , Shiyin Kang , Helen Meng

This work presents a framework based on feature disentanglement to learn speaker embeddings that are robust to environmental variations. Our framework utilises an auto-encoder as a disentangler, dividing the input speaker embedding into…

Sound · Computer Science 2024-06-21 KiHyun Nam , Hee-Soo Heo , Jee-weon Jung , Joon Son Chung

Voice conversion (VC) is a task that transforms voice from target audio to source without losing linguistic contents, it is challenging especially when source and target speakers are unseen during training (zero-shot VC). Previous…

Sound · Computer Science 2021-04-14 Shijun Wang , Damian Borth

Speaker identity is one of the important characteristics of human speech. In voice conversion, we change the speaker identity from one to another, while keeping the linguistic content unchanged. Voice conversion involves multiple speech…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-18 Berrak Sisman , Junichi Yamagishi , Simon King , Haizhou Li

The goal of accent conversion (AC) is to convert speech accents while preserving content and speaker identity. Previous methods either required reference utterances during inference, did not preserve speaker identity well, or used…

Sound · Computer Science 2024-10-08 Tuan Nam Nguyen , Ngoc Quan Pham , Alexander Waibel

Any-to-any voice conversion problem aims to convert voices for source and target speakers, which are out of the training data. Previous works wildly utilize the disentangle-based models. The disentangle-based model assumes the speech…

Sound · Computer Science 2022-02-23 Qiqi Wang , Xulong Zhang , Jianzong Wang , Ning Cheng , Jing Xiao

Speech signals encompass various information across multiple levels including content, speaker, and style. Disentanglement of these information, although challenging, is important for applications such as voice conversion. The contrastive…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-06 Yuying Xie , Michael Kuhlmann , Frederik Rautenberg , Zheng-Hua Tan , Reinhold Haeb-Umbach

Voice conversion (VC) consists of digitally altering the voice of an individual to manipulate part of its content, primarily its identity, while maintaining the rest unchanged. Research in neural VC has accomplished considerable…

Sound · Computer Science 2021-07-28 Laurent Benaroya , Nicolas Obin , Axel Roebel