English

Pushing the limits of raw waveform speaker recognition

Audio and Speech Processing 2022-03-30 v2 Artificial Intelligence

Abstract

In recent years, speaker recognition systems based on raw waveform inputs have received increasing attention. However, the performance of such systems are typically inferior to the state-of-the-art handcrafted feature-based counterparts, which demonstrate equal error rates under 1% on the popular VoxCeleb1 test set. This paper proposes a novel speaker recognition model based on raw waveform inputs. The model incorporates recent advances in machine learning and speaker verification, including the Res2Net backbone module and multi-layer feature aggregation. Our best model achieves an equal error rate of 0.89%, which is competitive with the state-of-the-art models based on handcrafted features, and outperforms the best model based on raw waveform inputs by a large margin. We also explore the application of the proposed model in the context of self-supervised learning framework. Our self-supervised model outperforms single phase-based existing works in this line of research. Finally, we show that self-supervised pre-training is effective for the semi-supervised scenario where we only have a small set of labelled training data, along with a larger set of unlabelled examples.

Keywords

Cite

@article{arxiv.2203.08488,
  title  = {Pushing the limits of raw waveform speaker recognition},
  author = {Jee-weon Jung and You Jin Kim and Hee-Soo Heo and Bong-Jin Lee and Youngki Kwon and Joon Son Chung},
  journal= {arXiv preprint arXiv:2203.08488},
  year   = {2022}
}

Comments

submitted to INTERSPEECH 2022 as a conference paper. 5 pages, 2 figures, 5 tables

R2 v1 2026-06-24T10:15:23.764Z