English

Multi-modal Emotion Estimation for in-the-wild Videos

Computer Vision and Pattern Recognition 2022-04-01 v4 Image and Video Processing

Abstract

In this paper, we briefly introduce our submission to the Valence-Arousal Estimation Challenge of the 3rd Affective Behavior Analysis in-the-wild (ABAW) competition. Our method utilizes the multi-modal information, i.e., the visual and audio information, and employs a temporal encoder to model the temporal context in the videos. Besides, a smooth processor is applied to get more reasonable predictions, and a model ensemble strategy is used to improve the performance of our proposed method. The experiment results show that our method achieves 65.55% ccc for valence and 70.88% ccc for arousal on the validation set of the Aff-Wild2 dataset, which prove the effectiveness of our proposed method.

Keywords

Cite

@article{arxiv.2203.13032,
  title  = {Multi-modal Emotion Estimation for in-the-wild Videos},
  author = {Liyu Meng and Yuchen Liu and Xiaolong Liu and Zhaopei Huang and Yuan Cheng and Meng Wang and Chuanhe Liu and Qin Jin},
  journal= {arXiv preprint arXiv:2203.13032},
  year   = {2022}
}
R2 v1 2026-06-24T10:24:36.568Z