English

Joint Multimodal Transformer for Emotion Recognition in the Wild

Computer Vision and Pattern Recognition 2024-04-23 v3 Machine Learning Sound Audio and Speech Processing

Abstract

Multimodal emotion recognition (MMER) systems typically outperform unimodal systems by leveraging the inter- and intra-modal relationships between, e.g., visual, textual, physiological, and auditory modalities. This paper proposes an MMER method that relies on a joint multimodal transformer (JMT) for fusion with key-based cross-attention. This framework can exploit the complementary nature of diverse modalities to improve predictive accuracy. Separate backbones capture intra-modal spatiotemporal dependencies within each modality over video sequences. Subsequently, our JMT fusion architecture integrates the individual modality embeddings, allowing the model to effectively capture inter- and intra-modal relationships. Extensive experiments on two challenging expression recognition tasks -- (1) dimensional emotion recognition on the Affwild2 dataset (with face and voice) and (2) pain estimation on the Biovid dataset (with face and biosensors) -- indicate that our JMT fusion can provide a cost-effective solution for MMER. Empirical results show that MMER systems with our proposed fusion allow us to outperform relevant baseline and state-of-the-art methods.

Keywords

Cite

@article{arxiv.2403.10488,
  title  = {Joint Multimodal Transformer for Emotion Recognition in the Wild},
  author = {Paul Waligora and Haseeb Aslam and Osama Zeeshan and Soufiane Belharbi and Alessandro Lameiras Koerich and Marco Pedersoli and Simon Bacon and Eric Granger},
  journal= {arXiv preprint arXiv:2403.10488},
  year   = {2024}
}

Comments

10 pages, 4 figures, 6 tables, CVPRw 2024

R2 v1 2026-06-28T15:22:03.635Z