Latent space representation for multi-target speaker detection and identification with a sparse dataset using Triplet neural networks
Abstract
We present an approach to tackle the speaker recognition problem using Triplet Neural Networks. Currently, the -vector representation with probabilistic linear discriminant analysis (PLDA) is the most commonly used technique to solve this problem, due to high classification accuracy with a relatively short computation time. In this paper, we explore a neural network approach, namely Triplet Neural Networks (TNNs), to built a latent space for different classifiers to solve the Multi-Target Speaker Detection and Identification Challenge Evaluation 2018 (MCE 2018) dataset. This training set contains -vectors from 3,631 speakers, with only 3 samples for each speaker, thus making speaker recognition a challenging task. When using the train and development set for training both the TNN and baseline model (i.e., similarity evaluation directly on the -vector representation), our proposed model outperforms the baseline by 23%. When reducing the training data to only using the train set, our method results in 309 confusions for the Multi-target speaker identification task, which is 46% better than the baseline model. These results show that the representational power of TNNs is especially evident when training on small datasets with few instances available per class.
Cite
@article{arxiv.1910.01463,
title = {Latent space representation for multi-target speaker detection and identification with a sparse dataset using Triplet neural networks},
author = {Kin Wai Cheuk and Balamurali B. T. and Gemma Roig and Dorien Herremans},
journal= {arXiv preprint arXiv:1910.01463},
year = {2019}
}
Comments
Accepted for ASRU 2019