English

Facial Affect Recognition based on Multi Architecture Encoder and Feature Fusion for the ABAW7 Challenge

Computer Vision and Pattern Recognition 2024-07-29 v2

Abstract

In this paper, we present our approach to addressing the challenges of the 7th ABAW competition. The competition comprises three sub-challenges: Valence Arousal (VA) estimation, Expression (Expr) classification, and Action Unit (AU) detection. To tackle these challenges, we employ state-of-the-art models to extract powerful visual features. Subsequently, a Transformer Encoder is utilized to integrate these features for the VA, Expr, and AU sub-challenges. To mitigate the impact of varying feature dimensions, we introduce an affine module to align the features to a common dimension. Overall, our results significantly outperform the baselines.

Keywords

Cite

@article{arxiv.2407.12258,
  title  = {Facial Affect Recognition based on Multi Architecture Encoder and Feature Fusion for the ABAW7 Challenge},
  author = {Kang Shen and Xuxiong Liu and Boyan Wang and Jun Yao and Xin Liu and Yujie Guan and Yu Wang and Gengchen Li and Xiao Sun},
  journal= {arXiv preprint arXiv:2407.12258},
  year   = {2024}
}
R2 v1 2026-06-28T17:43:57.996Z