English

Mitigating Modality and Language-Style Gaps for Zero-Shot Video Moment Retrieval

Computer Vision and Pattern Recognition 2026-07-21 v1

Abstract

Zero-shot video moment retrieval aims to overcome the limitations of traditional approaches that require large-scale datasets annotated with text and its relevant temporal spans. Despite advances in pre-trained vision-language models and multimodal large language models, existing ZMR methods still heavily depend on query-to-video content similarity, making them vulnerable to modality and language-style gaps. These gaps lead to unreliable span proposals and unstable moment retrieval results. To address this issue, we propose Self-Similarity-based Moment Proposal and Scoring that instead exploits intrinsic relationships within videos, enabling robust span generation and scoring. By deriving self-similarity only from the video content, we circumvent the noisy and mismatched patterns of query-frame or query-caption similarities, thereby mitigating both modality and language-style gaps. Furthermore, we introduce a query-aware MLLM-based reasoning stage to further sharpen alignment between text and video. Extensive experiments demonstrate that Self-SiMS achieves state-of-the-art performance across ZMR benchmarks.

Cite

@article{arxiv.2607.19027,
  title  = {Mitigating Modality and Language-Style Gaps for Zero-Shot Video Moment Retrieval},
  author = {Jihyun Lee and Cheol-Ho Cho and Woojin Jun and Woojin Jeong and Jae-Pil Heo},
  journal= {arXiv preprint arXiv:2607.19027},
  year   = {2026}
}

Comments

ECCV 2026 (* These authors contributed equally.)

R2 v1 2026-07-22T20:50:04.089Z