English

Sec2Sec Co-attention for Video-Based Apparent Affective Prediction

Multimedia 2024-08-28 v1

Abstract

Video-based apparent affect detection plays a crucial role in video understanding, as it encompasses various elements such as vision, audio, audio-visual interactions, and spatiotemporal information, which are essential for accurate video predictions. However, existing approaches often focus on extracting only a subset of these elements, resulting in the limited predictive capacity of their models. To address this limitation, we propose a novel LSTM-based network augmented with a Transformer co-attention mechanism for predicting apparent affect in videos. We demonstrate that our proposed Sec2Sec Co-attention Transformer surpasses multiple state-of-the-art methods in predicting apparent affect on two widely used datasets: LIRIS-ACCEDE and First Impressions. Notably, our model offers interpretability, allowing us to examine the contributions of different time points to the overall prediction. The implementation is available at: https://github.com/nestor-sun/sec2sec.

Keywords

Cite

@article{arxiv.2408.15209,
  title  = {Sec2Sec Co-attention for Video-Based Apparent Affective Prediction},
  author = {Mingwei Sun and Kunpeng Zhang},
  journal= {arXiv preprint arXiv:2408.15209},
  year   = {2024}
}

Comments

5 pages, 3 figures

R2 v1 2026-06-28T18:25:40.865Z