English

Focal Loss based Residual Convolutional Neural Network for Speech Emotion Recognition

Audio and Speech Processing 2025-04-16 v1 Artificial Intelligence Machine Learning Sound Machine Learning

Abstract

This paper proposes a Residual Convolutional Neural Network (ResNet) based on speech features and trained under Focal Loss to recognize emotion in speech. Speech features such as Spectrogram and Mel-frequency Cepstral Coefficients (MFCCs) have shown the ability to characterize emotion better than just plain text. Further Focal Loss, first used in One-Stage Object Detectors, has shown the ability to focus the training process more towards hard-examples and down-weight the loss assigned to well-classified examples, thus preventing the model from being overwhelmed by easily classifiable examples.

Keywords

Cite

@article{arxiv.1906.05682,
  title  = {Focal Loss based Residual Convolutional Neural Network for Speech Emotion Recognition},
  author = {Suraj Tripathi and Abhay Kumar and Abhiram Ramesh and Chirag Singh and Promod Yenigalla},
  journal= {arXiv preprint arXiv:1906.05682},
  year   = {2025}
}

Comments

Accepted in CICLing 2019

R2 v1 2026-06-23T09:52:45.123Z