Precise video retrieval requires multi-modal correlations to handle unseen vocabulary and scenes, becoming more complex for lengthy videos where models must perform effectively without prior training on a specific dataset. We introduce a unified framework that combines a visual matching stream and an aural matching stream with a unique subtitles-based video segmentation approach. Additionally, the aural stream includes a complementary audio-based two-stage retrieval mechanism that enhances performance on long-duration videos. Considering the complex nature of retrieval from lengthy videos and its corresponding evaluation, we introduce a new retrieval evaluation method specifically designed for long-video retrieval to support further research. We conducted experiments on the YouCook2 benchmark, showing promising retrieval performance.
@article{arxiv.2504.04572,
title = {Multimodal Lengthy Videos Retrieval Framework and Evaluation Metric},
author = {Mohamed Eltahir and Osamah Sarraj and Mohammed Bremoo and Mohammed Khurd and Abdulrahman Alfrihidi and Taha Alshatiri and Mohammad Almatrafi and Tanveer Hussain},
journal= {arXiv preprint arXiv:2504.04572},
year = {2025}
}