Beijing ZKJ-NPU Speaker Verification System for VoxCeleb Speaker Recognition Challenge 2021
Abstract
In this report, we describe the Beijing ZKJ-NPU team submission to the VoxCeleb Speaker Recognition Challenge 2021 (VoxSRC-21). We participated in the fully supervised speaker verification track 1 and track 2. In the challenge, we explored various kinds of advanced neural network structures with different pooling layers and objective loss functions. In addition, we introduced the ResNet-DTCF, CoAtNet and PyConv networks to advance the performance of CNN-based speaker embedding model. Moreover, we applied embedding normalization and score normalization at the evaluation stage. By fusing 11 and 14 systems, our final best performances (minDCF/EER) on the evaluation trails are 0.1205/2.8160% and 0.1175/2.8400% respectively for track 1 and 2. With our submission, we came to the second place in the challenge for both tracks.
Keywords
Cite
@article{arxiv.2109.03568,
title = {Beijing ZKJ-NPU Speaker Verification System for VoxCeleb Speaker Recognition Challenge 2021},
author = {Li Zhang and Huan Zhao and Qinling Meng and Yanli Chen and Min Liu and Lei Xie},
journal= {arXiv preprint arXiv:2109.03568},
year = {2021}
}