English

Closed-Loop Open-Vocabulary Mobile Manipulation with GPT-4V

Robotics 2025-03-10 v2 Artificial Intelligence Computer Vision and Pattern Recognition Machine Learning

Abstract

Autonomous robot navigation and manipulation in open environments require reasoning and replanning with closed-loop feedback. In this work, we present COME-robot, the first closed-loop robotic system utilizing the GPT-4V vision-language foundation model for open-ended reasoning and adaptive planning in real-world scenarios.COME-robot incorporates two key innovative modules: (i) a multi-level open-vocabulary perception and situated reasoning module that enables effective exploration of the 3D environment and target object identification using commonsense knowledge and situated information, and (ii) an iterative closed-loop feedback and restoration mechanism that verifies task feasibility, monitors execution success, and traces failure causes across different modules for robust failure recovery. Through comprehensive experiments involving 8 challenging real-world mobile and tabletop manipulation tasks, COME-robot demonstrates a significant improvement in task success rate (~35%) compared to state-of-the-art methods. We further conduct comprehensive analyses to elucidate how COME-robot's design facilitates failure recovery, free-form instruction following, and long-horizon task planning.

Keywords

Cite

@article{arxiv.2404.10220,
  title  = {Closed-Loop Open-Vocabulary Mobile Manipulation with GPT-4V},
  author = {Peiyuan Zhi and Zhiyuan Zhang and Yu Zhao and Muzhi Han and Zeyu Zhang and Zhitian Li and Ziyuan Jiao and Baoxiong Jia and Siyuan Huang},
  journal= {arXiv preprint arXiv:2404.10220},
  year   = {2025}
}

Comments

6 pages, Accepted at 2025 IEEE ICRA, website: https://come-robot.github.io/

R2 v1 2026-06-28T15:55:17.904Z