ObjChangeVR:来自VR环境中连续第三人称视图的对象状态变化推理
计算机视觉与模式识别
2026-03-10 v1 人工智能
摘要
多模态大语言模型(MLLM)的最新进展为虚拟现实(VR)中的自然语言场景变更查询提供了有前景的途径。先前将MLLM应用于对象状态理解的工作主要关注应用于对象状态理解的第三人称视频,这些视频捕获摄像机佩戴者与对象的交互。然而,对象状态更改可能发生在背景中且无需直接用户交互, lacks explicit motion cues and making them difficult to detect. 此外,尚无基准用于评估此具有挑战性的情景。为解决这些挑战,我们引入ObjChangeVR-Dataset, specifically for benchmarking the question-answering task of object state change. 我们还提出了ObjChangeVR, a framework that combines viewpoint-aware and temporal-based retrieval to identify relevant frames, along with cross-view reasoning that reconciles inconsistent evidence from multiple viewpoints. 广泛的实验表明,ObjChangeVR在多个MLLMs上显著优于基线方法。
引用
@article{arxiv.2603.06648,
title = {ObjChangeVR: Object State Change Reasoning from Continuous Egocentric Views in VR Environments},
author = {Shiyi Ding and Shaoen Wu and Ying Chen},
journal= {arXiv preprint arXiv:2603.06648},
year = {2026}
}
备注
European Chapter of the Association for Computational Linguistics (EACL) 2026 Main