English

Zero-Shot Video Moment Retrieval from Frozen Vision-Language Models

Computer Vision and Pattern Recognition 2023-09-06 v1

Abstract

Accurate video moment retrieval (VMR) requires universal visual-textual correlations that can handle unknown vocabulary and unseen scenes. However, the learned correlations are likely either biased when derived from a limited amount of moment-text data which is hard to scale up because of the prohibitive annotation cost (fully-supervised), or unreliable when only the video-text pairwise relationships are available without fine-grained temporal annotations (weakly-supervised). Recently, the vision-language models (VLM) demonstrate a new transfer learning paradigm to benefit different vision tasks through the universal visual-textual correlations derived from large-scale vision-language pairwise web data, which has also shown benefits to VMR by fine-tuning in the target domains. In this work, we propose a zero-shot method for adapting generalisable visual-textual priors from arbitrary VLM to facilitate moment-text alignment, without the need for accessing the VMR data. To this end, we devise a conditional feature refinement module to generate boundary-aware visual features conditioned on text queries to enable better moment boundary understanding. Additionally, we design a bottom-up proposal generation strategy that mitigates the impact of domain discrepancies and breaks down complex-query retrieval tasks into individual action retrievals, thereby maximizing the benefits of VLM. Extensive experiments conducted on three VMR benchmark datasets demonstrate the notable performance advantages of our zero-shot algorithm, especially in the novel-word and novel-location out-of-distribution setups.

Keywords

Cite

@article{arxiv.2309.00661,
  title  = {Zero-Shot Video Moment Retrieval from Frozen Vision-Language Models},
  author = {Dezhao Luo and Jiabo Huang and Shaogang Gong and Hailin Jin and Yang Liu},
  journal= {arXiv preprint arXiv:2309.00661},
  year   = {2023}
}

Comments

Accepted by WACV 2024