English

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding

Computer Vision and Pattern Recognition 2025-07-08 v1 Artificial Intelligence

Abstract

Fine-grained video classification requires understanding complex spatio-temporal and semantic cues that often exceed the capacity of a single modality. In this paper, we propose a multimodal framework that fuses video, image, and text representations using GRU-based sequence encoders and cross-modal attention mechanisms. The model is trained using a combination of classification or regression loss, depending on the task, and is further regularized through feature-level augmentation and autoencoding techniques. To evaluate the generality of our framework, we conduct experiments on two challenging benchmarks: the DVD dataset for real-world violence detection and the Aff-Wild2 dataset for valence-arousal estimation. Our results demonstrate that the proposed fusion strategy significantly outperforms unimodal baselines, with cross-attention and feature augmentation contributing notably to robustness and performance.

Keywords

Cite

@article{arxiv.2507.03531,
  title  = {Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding},
  author = {Namho Kim and Junhwa Kim},
  journal= {arXiv preprint arXiv:2507.03531},
  year   = {2025}
}
R2 v1 2026-07-01T03:46:42.652Z