English

SpeakerNet: 1D Depth-wise Separable Convolutional Network for Text-Independent Speaker Recognition and Verification

Audio and Speech Processing 2020-10-27 v1

Abstract

We propose SpeakerNet - a new neural architecture for speaker recognition and speaker verification tasks. It is composed of residual blocks with 1D depth-wise separable convolutions, batch-normalization, and ReLU layers. This architecture uses x-vector based statistics pooling layer to map variable-length utterances to a fixed-length embedding (q-vector). SpeakerNet-M is a simple lightweight model with just 5M parameters. It doesn't use voice activity detection (VAD) and achieves close to state-of-the-art performance scoring an Equal Error Rate (EER) of 2.10% on the VoxCeleb1 cleaned and 2.29% on the VoxCeleb1 trial files.

Keywords

Cite

@article{arxiv.2010.12653,
  title  = {SpeakerNet: 1D Depth-wise Separable Convolutional Network for Text-Independent Speaker Recognition and Verification},
  author = {Nithin Rao Koluguri and Jason Li and Vitaly Lavrukhin and Boris Ginsburg},
  journal= {arXiv preprint arXiv:2010.12653},
  year   = {2020}
}

Comments

Preprint, submitted to ICASSP 2021

R2 v1 2026-06-23T19:36:18.337Z