基于中心-外周视野的多模态模型用于空间推理
计算机视觉与模式识别
2025-12-10 v1
摘要
我们提出了一种中心-外周视野启发的框架(CVP),一种简单而有效的多模态模型用于空间推理, draw inspiration from 两种类型的人类视野——中心视野和外周视野。现有方法主要依赖于非结构化表示,如点云、体素或 patch features, 并通过坐标嵌入隐式地注入场景上下文。然而,这往往导致受限的空间推理能力 due to the lack of explicit, high-level structural understanding。为 address this limitation,我们将两种互补组件引入 Large Multimodal Model-based architecture: target-affinity token,类似于中心视野, 指向 query-relevant objects 的注意力;以及 allocentric grid,类似于外周视野, 捕获全局场景上下文和空间布局。这些组件共同工作,以 enable 对复杂 3D 环境的结构化、上下文感知的理解。实验表明,CVP 在 various 3D 场景理解基准上实现了 state-of-the-art performance。
引用
@article{arxiv.2512.08135,
title = {CVP: Central-Peripheral Vision-Inspired Multimodal Model for Spatial Reasoning},
author = {Zeyuan Chen and Xiang Zhang and Haiyang Xu and Jianwen Xie and Zhuowen Tu},
journal= {arXiv preprint arXiv:2512.08135},
year = {2025}
}
备注
Accepted to WACV 2026