English

$M^3$T: Multi-Modal Continuous Valence-Arousal Estimation in the Wild

Computer Vision and Pattern Recognition 2020-02-10 v1 Sound Audio and Speech Processing Image and Video Processing

Abstract

This report describes a multi-modal multi-task (M3M^3T) approach underlying our submission to the valence-arousal estimation track of the Affective Behavior Analysis in-the-wild (ABAW) Challenge, held in conjunction with the IEEE International Conference on Automatic Face and Gesture Recognition (FG) 2020. In the proposed M3M^3T framework, we fuse both visual features from videos and acoustic features from the audio tracks to estimate the valence and arousal. The spatio-temporal visual features are extracted with a 3D convolutional network and a bidirectional recurrent neural network. Considering the correlations between valence / arousal, emotions, and facial actions, we also explores mechanisms to benefit from other tasks. We evaluated the M3M^3T framework on the validation set provided by ABAW and it significantly outperforms the baseline method.

Keywords

Cite

@article{arxiv.2002.02957,
  title  = {$M^3$T: Multi-Modal Continuous Valence-Arousal Estimation in the Wild},
  author = {Yuan-Hang Zhang and Rulin Huang and Jiabei Zeng and Shiguang Shan and Xilin Chen},
  journal= {arXiv preprint arXiv:2002.02957},
  year   = {2020}
}

Comments

6 pages, technical report; submission to ABAW Challenge at FG 2020

R2 v1 2026-06-23T13:34:39.192Z