English

CASA-Based Speaker Identification Using Cascaded GMM-CNN Classifier in Noisy and Emotional Talking Conditions

Sound 2021-02-12 v1 Artificial Intelligence Audio and Speech Processing

Abstract

This work aims at intensifying text-independent speaker identification performance in real application situations such as noisy and emotional talking conditions. This is achieved by incorporating two different modules: a Computational Auditory Scene Analysis CASA based pre-processing module for noise reduction and cascaded Gaussian Mixture Model Convolutional Neural Network GMM-CNN classifier for speaker identification followed by emotion recognition. This research proposes and evaluates a novel algorithm to improve the accuracy of speaker identification in emotional and highly-noise susceptible conditions. Experiments demonstrate that the proposed model yields promising results in comparison with other classifiers when Speech Under Simulated and Actual Stress SUSAS database, Emirati Speech Database ESD, the Ryerson Audio-Visual Database of Emotional Speech and Song RAVDESS database and the Fluent Speech Commands database are used in a noisy environment.

Keywords

Cite

@article{arxiv.2102.05894,
  title  = {CASA-Based Speaker Identification Using Cascaded GMM-CNN Classifier in Noisy and Emotional Talking Conditions},
  author = {Ali Bou Nassif and Ismail Shahin and Shibani Hamsa and Nawel Nemmour and Keikichi Hirose},
  journal= {arXiv preprint arXiv:2102.05894},
  year   = {2021}
}

Comments

Published in Applied Soft Computing journal