English

GIST-AiTeR System for the Diarization Task of the 2022 VoxCeleb Speaker Recognition Challenge

Audio and Speech Processing 2022-10-07 v4

Abstract

This report describes the submission system of the GIST-AiTeR team at the 2022 VoxCeleb Speaker Recognition Challenge (VoxSRC) Track 4. Our system mainly includes speech enhancement, voice activity detection , multi-scaled speaker embedding, probabilistic linear discriminant analysis-based speaker clustering, and overlapped speech detection models. We first construct four different diarization systems according to different model combinations with the best experimental efforts. Our final submission is an ensemble system of all the four systems and achieves a diarization error rate of 5.12% on the challenge evaluation set, ranked third at the diarization track of the challenge.

Keywords

Cite

@article{arxiv.2209.10357,
  title  = {GIST-AiTeR System for the Diarization Task of the 2022 VoxCeleb Speaker Recognition Challenge},
  author = {Dongkeon Park and Yechan Yu and Kyeong Wan Park and Ji Won Kim and Hong Kook Kim},
  journal= {arXiv preprint arXiv:2209.10357},
  year   = {2022}
}

Comments

2022 VoxSRC Track4