English

The DKU-Tencent System for the VoxCeleb Speaker Recognition Challenge 2022

Sound 2022-10-12 v1 Audio and Speech Processing

Abstract

This paper is the system description of the DKU-Tencent System for the VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC22). In this challenge, we focus on track1 and track3. For track1, multiple backbone networks are adopted to extract frame-level features. Since track1 focus on the cross-age scenarios, we adopt the cross-age trials and perform QMF to calibrate score. The magnitude-based quality measures achieve a large improvement. For track3, the semi-supervised domain adaptation task, the pseudo label method is adopted to make domain adaptation. Considering the noise labels in clustering, the ArcFace is replaced by Sub-center ArcFace. The final submission achieves 0.107 mDCF in task1 and 7.135% EER in task3.

Keywords

Cite

@article{arxiv.2210.05092,
  title  = {The DKU-Tencent System for the VoxCeleb Speaker Recognition Challenge 2022},
  author = {Xiaoyi Qin and Na Li and Yuke Lin and Yiwei Ding and Chao Weng and Dan Su and Ming Li},
  journal= {arXiv preprint arXiv:2210.05092},
  year   = {2022}
}