English

Solution to the 10th ABAW Expression Recognition Challenge: A Robust Multimodal Framework with Safe Cross-Attention and Modality Dropout

Computer Vision and Pattern Recognition 2026-03-10 v1 Artificial Intelligence

Abstract

Emotion recognition in real-world environments is hindered by partial occlusions, missing modalities, and severe class imbalance. To address these issues, particularly for the Affective Behavior Analysis in-the-wild (ABAW) Expression challenge, we propose a multimodal framework that dynamically fuses visual and audio representations. Our approach uses a dual-branch Transformer architecture featuring a safe cross-attention mechanism and a modality dropout strategy. This design allows the network to rely on audio-based predictions when visual cues are absent. To mitigate the long-tail distribution of the Aff-Wild2 dataset, we apply focal loss optimization, combined with a sliding-window soft voting strategy to capture dynamic emotional transitions and reduce frame-level classification jitter. Experiments demonstrate that our framework effectively handles missing modalities and complex spatiotemporal dependencies, achieving an accuracy of 60.79% and an F1-score of 0.5029 on the Aff-Wild2 validation set.

Keywords

Cite

@article{arxiv.2603.08034,
  title  = {Solution to the 10th ABAW Expression Recognition Challenge: A Robust Multimodal Framework with Safe Cross-Attention and Modality Dropout},
  author = {Jun Yu and Naixiang Zheng and Guoyuan Wang and Yunxiang Zhang and Lingsi Zhu and Jiaen Liang and Wei Huang and Shengping Liu},
  journal= {arXiv preprint arXiv:2603.08034},
  year   = {2026}
}
R2 v1 2026-07-01T11:09:46.143Z