English

Latent space representation for multi-target speaker detection and identification with a sparse dataset using Triplet neural networks

Sound 2019-10-07 v2 Machine Learning Audio and Speech Processing

Abstract

We present an approach to tackle the speaker recognition problem using Triplet Neural Networks. Currently, the ii-vector representation with probabilistic linear discriminant analysis (PLDA) is the most commonly used technique to solve this problem, due to high classification accuracy with a relatively short computation time. In this paper, we explore a neural network approach, namely Triplet Neural Networks (TNNs), to built a latent space for different classifiers to solve the Multi-Target Speaker Detection and Identification Challenge Evaluation 2018 (MCE 2018) dataset. This training set contains ii-vectors from 3,631 speakers, with only 3 samples for each speaker, thus making speaker recognition a challenging task. When using the train and development set for training both the TNN and baseline model (i.e., similarity evaluation directly on the ii-vector representation), our proposed model outperforms the baseline by 23%. When reducing the training data to only using the train set, our method results in 309 confusions for the Multi-target speaker identification task, which is 46% better than the baseline model. These results show that the representational power of TNNs is especially evident when training on small datasets with few instances available per class.

Keywords

Cite

@article{arxiv.1910.01463,
  title  = {Latent space representation for multi-target speaker detection and identification with a sparse dataset using Triplet neural networks},
  author = {Kin Wai Cheuk and Balamurali B. T. and Gemma Roig and Dorien Herremans},
  journal= {arXiv preprint arXiv:1910.01463},
  year   = {2019}
}

Comments

Accepted for ASRU 2019

R2 v1 2026-06-23T11:33:43.199Z