English

Emotion-Disentangled Embedding Alignment for Noise-Robust and Cross-Corpus Speech Emotion Recognition

Sound 2025-10-13 v1 Artificial Intelligence Human-Computer Interaction Machine Learning Audio and Speech Processing

Abstract

Effectiveness of speech emotion recognition in real-world scenarios is often hindered by noisy environments and variability across datasets. This paper introduces a two-step approach to enhance the robustness and generalization of speech emotion recognition models through improved representation learning. First, our model employs EDRL (Emotion-Disentangled Representation Learning) to extract class-specific discriminative features while preserving shared similarities across emotion categories. Next, MEA (Multiblock Embedding Alignment) refines these representations by projecting them into a joint discriminative latent subspace that maximizes covariance with the original speech input. The learned EDRL-MEA embeddings are subsequently used to train an emotion classifier using clean samples from publicly available datasets, and are evaluated on unseen noisy and cross-corpus speech samples. Improved performance under these challenging conditions demonstrates the effectiveness of the proposed method.

Keywords

Cite

@article{arxiv.2510.09072,
  title  = {Emotion-Disentangled Embedding Alignment for Noise-Robust and Cross-Corpus Speech Emotion Recognition},
  author = {Upasana Tiwari and Rupayan Chakraborty and Sunil Kumar Kopparapu},
  journal= {arXiv preprint arXiv:2510.09072},
  year   = {2025}
}

Comments

13 pages, 1 figure

R2 v1 2026-07-01T06:28:47.850Z