English

Noro: Noise-Robust One-shot Voice Conversion with Hidden Speaker Representation Learning

Sound 2025-08-29 v2 Computation and Language Audio and Speech Processing

Abstract

The effectiveness of one-shot voice conversion (VC) decreases in real-world scenarios where reference speeches, which are often sourced from the internet, contain various disturbances like background noise. To address this issue, we introduce Noro, a noise-robust one-shot VC system. Noro features innovative components tailored for VC using noisy reference speeches, including a dual-branch reference encoding module and a noise-agnostic contrastive speaker loss. Experimental results demonstrate that Noro outperforms our baseline system in both clean and noisy scenarios, highlighting its efficacy for real-world applications. Additionally, we investigate the hidden speaker representation capabilities of our baseline system by repurposing its reference encoder as a speaker encoder. The results show that it is competitive with several advanced self-supervised learning models for speaker representation under the SUPERB settings, highlighting the potential for advancing speaker representation learning through one-shot VC tasks.

Keywords

Cite

@article{arxiv.2411.19770,
  title  = {Noro: Noise-Robust One-shot Voice Conversion with Hidden Speaker Representation Learning},
  author = {Haorui He and Yuchen Song and Yuancheng Wang and Haoyang Li and Xueyao Zhang and Li Wang and Gongping Huang and Eng Siong Chng and Zhizheng Wu},
  journal= {arXiv preprint arXiv:2411.19770},
  year   = {2025}
}

Comments

Accepted by APSIPA ASC 2025

R2 v1 2026-06-28T20:16:54.833Z