面向整体事件感知与匹配的 E.M.Ground:一种时空 grounding Vid-LLM
计算机视觉与模式识别
2026-02-06 v1
摘要
尽管近期在视频大语言模型(Vid-LLMs)方面取得了进展,但 Temporal Video Grounding(TVG)——即精确定位查询事件对应的时间段——仍然是一个重大挑战。现有方法常通过比较帧特征与两个独立 token 来匹配起始帧和结束帧,高度依赖精确时间戳。然而,这种方法未能捕捉事件的语义连续性和完整性,导致歧义。为此,我们提出 E.M.Ground,一种聚焦于整体和连贯事件感知的新型 Vid-LLM。E.M.Ground 引入三项关键创新:(i) 一种特殊的 <event> token,用于聚合查询事件所有帧的信息,保留语义连续性以实现准确的事件匹配;(ii) Savitzky-Golay 滤波以减少 token 到帧相似性在时间戳上的噪声,提高预测准确度;(iii) 多层次帧特征聚合以增强匹配可靠性和时空理解,弥补压缩导致的信息损失。大量实验表明,E.M.Ground 在基准数据集上 consistently 超过现有最先进 Vid-LLM,获得显著的性能提升。
引用
@article{arxiv.2602.05215,
title = {E.M.Ground: A Temporal Grounding Vid-LLM with Holistic Event Perception and Matching},
author = {Jiahao Nie and Wenbin An and Gongjie Zhang and Yicheng Xu and Yap-Peng Tan and Alex C. Kot and Shijian Lu},
journal= {arXiv preprint arXiv:2602.05215},
year = {2026}
}