English

HourVideo: 1-Hour Video-Language Understanding

Computer Vision and Pattern Recognition 2024-11-08 v1 Artificial Intelligence Machine Learning

Abstract

We present HourVideo, a benchmark dataset for hour-long video-language understanding. Our dataset consists of a novel task suite comprising summarization, perception (recall, tracking), visual reasoning (spatial, temporal, predictive, causal, counterfactual), and navigation (room-to-room, object retrieval) tasks. HourVideo includes 500 manually curated egocentric videos from the Ego4D dataset, spanning durations of 20 to 120 minutes, and features 12,976 high-quality, five-way multiple-choice questions. Benchmarking results reveal that multimodal models, including GPT-4 and LLaVA-NeXT, achieve marginal improvements over random chance. In stark contrast, human experts significantly outperform the state-of-the-art long-context multimodal model, Gemini Pro 1.5 (85.0% vs. 37.3%), highlighting a substantial gap in multimodal capabilities. Our benchmark, evaluation toolkit, prompts, and documentation are available at https://hourvideo.stanford.edu

Keywords

Cite

@article{arxiv.2411.04998,
  title  = {HourVideo: 1-Hour Video-Language Understanding},
  author = {Keshigeyan Chandrasegaran and Agrim Gupta and Lea M. Hadzic and Taran Kota and Jimming He and Cristóbal Eyzaguirre and Zane Durante and Manling Li and Jiajun Wu and Li Fei-Fei},
  journal= {arXiv preprint arXiv:2411.04998},
  year   = {2024}
}

Comments

NeurIPS 2024 Datasets and Benchmarks Track; 28 pages

R2 v1 2026-06-28T19:52:07.198Z