English

SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion

Sound 2025-10-13 v1 Audio and Speech Processing

Abstract

Voice Conversion (VC) aims to modify a speaker's timbre while preserving linguistic content. While recent VC models achieve strong performance, most struggle in real-time streaming scenarios due to high latency, dependence on ASR modules, or complex speaker disentanglement, which often results in timbre leakage or degraded naturalness. We present SynthVC, a streaming end-to-end VC framework that directly learns speaker timbre transformation from synthetic parallel data generated by a pre-trained zero-shot VC model. This design eliminates the need for explicit content-speaker separation or recognition modules. Built upon a neural audio codec architecture, SynthVC supports low-latency streaming inference with high output fidelity. Experimental results show that SynthVC outperforms baseline streaming VC systems in both naturalness and speaker similarity, achieving an end-to-end latency of just 77.1 ms.

Keywords

Cite

@article{arxiv.2510.09245,
  title  = {SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion},
  author = {Zhao Guo and Ziqian Ning and Guobin Ma and Lei Xie},
  journal= {arXiv preprint arXiv:2510.09245},
  year   = {2025}
}

Comments

Accepted by NCMMSC2025

R2 v1 2026-07-01T06:29:08.749Z