English

Language-Invariant Multilingual Speaker Verification for the TidyVoice 2026 Challenge

Audio and Speech Processing 2026-03-10 v1

Abstract

Multilingual speaker verification (SV) remains challenging due to limited cross-lingual data and language-dependent information in speaker embeddings. This paper presents a language-invariant multilingual SV system for the TidyVoice 2026 Challenge. We adopt the multilingual self-supervised w2v-BERT 2.0 model as the backbone, enhanced with Layer Adapters and Multi-scale Feature Aggregation to better exploit multi-layer representations. A language-adversarial training strategy with a Gradient Reversal Layer is applied to promote language-invariant speaker embeddings. Moreover, a multilingual zero-shot text-to-speech system is used to synthesize speech in multiple languages, improving language diversity. Experimental results demonstrate that fine-tuning the large-scale pretrained model yields competitive performance, while language-adversarial training further enhances robustness. In addition, synthetic speech augmentation provides additional gains under limited training data conditions. Source code is available at https://github.com/ZXHY-82/LI-MSV-TidyVoice2026.

Keywords

Cite

@article{arxiv.2603.08092,
  title  = {Language-Invariant Multilingual Speaker Verification for the TidyVoice 2026 Challenge},
  author = {Ze Li and Xiaoxiao Miao and Juan Liu and Ming Li},
  journal= {arXiv preprint arXiv:2603.08092},
  year   = {2026}
}

Comments

submitted to Interspeech 2026

R2 v1 2026-07-01T11:09:51.486Z