中文

从预测时序感知到证据驱动推理:视频时序 grounding 的 Foresee-to-Ground 框架

计算机视觉与模式识别 2026-05-22 v1

摘要

当前的 Video-LLM 方法用于视频时序 grounding(VTG)通常依赖于从非结构化视觉 token 流中直接生成时间戳,常导致数值脆弱且边界不一致。为此,我们提出 Foresee-to-Ground(F2G),一种将 VTG 重新表述为可验证的 Identify-then-Measure 问题的框架。F2G 集成了预测时序感知与证据驱动推理:它学习边界敏感的时序表示,以构建视频范围的证据池,筛选出候选事件段,并将这些段作为可引用的证据单元暴露给 LLM,将边界预测绑定到显式事件假设上。通过解耦事件识别与精确边界测量,F2G 稳定了 grounding 并使预测可验证。广泛的实验表明,F2G 在各种基准上持续提高 grounding 准确率,能够在不同的 Video-LLM 背板之间实现稳健的迁移,同时保持一般视频理解能力。

关键词

引用

@article{arxiv.2605.21973,
  title  = {Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding},
  author = {Zelin Zheng and Xinyan Liu and Ruixin Li and Antoni B. Chan and Guorong Li and Qingming Huang and Laiyun Qing},
  journal= {arXiv preprint arXiv:2605.21973},
  year   = {2026}
}

备注

Accepted by ICML 2026