VideoRefer Suite:用于视频大语言模型中空间-时间对象理解的进阶方案
计算机视觉与模式识别
2025-03-26 v3 人工智能
机器学习
摘要
视频大语言模型(Video LLMs)最近在通用视频理解方面展现出惊人的能力。然而,它们主要关注整体理解,难以捕捉细粒度的空间和时间细节。此外,缺乏高质量的对象级视频指令数据以及全面的基准测试进一步阻碍了它们的发展。为解决这些挑战,我们引入VideoRefer Suite 以赋能Video LLM进行更细粒度的空间-时间视频理解,即实现对视频中任何对象的感知与推理。具体而言,我们在三个关键方面全面开发VideoRefer Suite:数据集、模型和基准测试。首先,我们引入一个多智能体数据引擎,精心筛选出大规模高质量的对象级视频指令数据集,称为VideoRefer-700K。接着,我们提出VideoRefer模型, equipped with a versatile spatial-temporal object encoder to capture precise regional and sequential representations. 最后,我们精心创建VideoRefer-Bench 以全面评估Video LLM的空间-时间理解能力,跨越various aspects。大量实验和分析表明,我们的VideoRefer模型不仅在视频指涉基准测试中取得优异成绩,也促进了通用视频理解能力的发展。
引用
@article{arxiv.2501.00599,
title = {VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM},
author = {Yuqian Yuan and Hang Zhang and Wentong Li and Zesen Cheng and Boqiang Zhang and Long Li and Xin Li and Deli Zhao and Wenqiao Zhang and Yueting Zhuang and Jianke Zhu and Lidong Bing},
journal= {arXiv preprint arXiv:2501.00599},
year = {2025}
}
备注
17 pages, 14 figures, technical report