English

Multimodal Lengthy Videos Retrieval Framework and Evaluation Metric

Computer Vision and Pattern Recognition 2025-04-08 v1

Abstract

Precise video retrieval requires multi-modal correlations to handle unseen vocabulary and scenes, becoming more complex for lengthy videos where models must perform effectively without prior training on a specific dataset. We introduce a unified framework that combines a visual matching stream and an aural matching stream with a unique subtitles-based video segmentation approach. Additionally, the aural stream includes a complementary audio-based two-stage retrieval mechanism that enhances performance on long-duration videos. Considering the complex nature of retrieval from lengthy videos and its corresponding evaluation, we introduce a new retrieval evaluation method specifically designed for long-video retrieval to support further research. We conducted experiments on the YouCook2 benchmark, showing promising retrieval performance.

Keywords

Cite

@article{arxiv.2504.04572,
  title  = {Multimodal Lengthy Videos Retrieval Framework and Evaluation Metric},
  author = {Mohamed Eltahir and Osamah Sarraj and Mohammed Bremoo and Mohammed Khurd and Abdulrahman Alfrihidi and Taha Alshatiri and Mohammad Almatrafi and Tanveer Hussain},
  journal= {arXiv preprint arXiv:2504.04572},
  year   = {2025}
}
R2 v1 2026-06-28T22:48:41.901Z