English

Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning

Computer Vision and Pattern Recognition 2025-07-25 v1 Artificial Intelligence

Abstract

Video Temporal Grounding (VTG) aims to localize relevant temporal segments in videos given natural language queries. Despite recent progress with large vision-language models (LVLMs) and instruction-tuning, existing approaches often suffer from limited temporal awareness and poor generalization. In this work, we introduce a two-stage training framework that integrates supervised fine-tuning with reinforcement learning (RL) to improve both the accuracy and robustness of VTG models. Our approach first leverages high-quality curated cold start data for SFT initialization, followed by difficulty-controlled RL to further enhance temporal localization and reasoning abilities. Comprehensive experiments on multiple VTG benchmarks demonstrate that our method consistently outperforms existing models, particularly in challenging and open-domain scenarios. We conduct an in-depth analysis of training strategies and dataset curation, highlighting the importance of both high-quality cold start data and difficulty-controlled RL. To facilitate further research and industrial adoption, we release all intermediate datasets, models, and code to the community.

Keywords

Cite

@article{arxiv.2507.18100,
  title  = {Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning},
  author = {Ruizhe Chen and Zhiting Fan and Tianze Luo and Heqing Zou and Zhaopeng Feng and Guiyang Xie and Hansheng Zhang and Zhuochen Wang and Zuozhu Liu and Huaijian Zhang},
  journal= {arXiv preprint arXiv:2507.18100},
  year   = {2025}
}
R2 v1 2026-07-01T04:16:26.882Z