English

Point What You Mean: Visually Grounded Instruction Policy

Computer Vision and Pattern Recognition 2026-03-25 v2 Robotics

Abstract

Vision-Language-Action (VLA) models align vision and language with embodied control, but their object referring ability remains limited when relying solely on text prompt, especially in cluttered or out-of-distribution (OOD) scenes. In this study, we introduce the Point-VLA, a plug-and-play policy that augments language instructions with explicit visual cues (e.g., bounding boxes) to resolve referential ambiguity and enable precise object-level grounding. To efficiently scale visually grounded datasets, we further develop an automatic data annotation pipeline requiring minimal human effort. We evaluate Point-VLA on diverse real-world referring tasks and observe consistently stronger performance than text-only instruction VLAs, particularly in cluttered or unseen-object scenarios, with robust generalization. These results demonstrate that Point-VLA effectively resolves object referring ambiguity through pixel-level visual grounding, achieving more generalizable embodied control.

Keywords

Cite

@article{arxiv.2512.18933,
  title  = {Point What You Mean: Visually Grounded Instruction Policy},
  author = {Hang Yu and Juntu Zhao and Yufeng Liu and Kaiyu Li and Cheng Ma and Di Zhang and Yingdong Hu and Guang Chen and Junyuan Xie and Junliang Guo and Junqiao Zhao and Yang Gao},
  journal= {arXiv preprint arXiv:2512.18933},
  year   = {2026}
}
R2 v1 2026-07-01T08:35:56.717Z