The DKU-Tencent System for the VoxCeleb Speaker Recognition Challenge 2022
Sound
2022-10-12 v1 Audio and Speech Processing
Abstract
This paper is the system description of the DKU-Tencent System for the VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC22). In this challenge, we focus on track1 and track3. For track1, multiple backbone networks are adopted to extract frame-level features. Since track1 focus on the cross-age scenarios, we adopt the cross-age trials and perform QMF to calibrate score. The magnitude-based quality measures achieve a large improvement. For track3, the semi-supervised domain adaptation task, the pseudo label method is adopted to make domain adaptation. Considering the noise labels in clustering, the ArcFace is replaced by Sub-center ArcFace. The final submission achieves 0.107 mDCF in task1 and 7.135% EER in task3.
Keywords
Cite
@article{arxiv.2210.05092,
title = {The DKU-Tencent System for the VoxCeleb Speaker Recognition Challenge 2022},
author = {Xiaoyi Qin and Na Li and Yuke Lin and Yiwei Ding and Chao Weng and Dan Su and Ming Li},
journal= {arXiv preprint arXiv:2210.05092},
year = {2022}
}