English

PathPainter: Transferring the Generalization Ability of Image Generation Models to Embodied Navigation

Robotics 2026-05-11 v1

Abstract

Bird's-eye-view (BEV) images have been widely demonstrated to provide valuable prior information for navigation. Given the global information provided by such views, two key challenges remain: how to fully exploit this information and how to reliably use it during execution. In this paper, we propose a navigation system that uses BEV images as global priors and is designed for ground and near-ground robotic platforms. The system employs an image generation model to interpret human intent from natural language, identify the target destination, and generate traversability masks. During execution, we introduce cross-view localization to align the robot's odometry with the BEV map and mitigate long-term drift in conventional odometry. We conduct extensive benchmark experiments to evaluate the proposed method and further validate it on a UAV platform. Using only a conventional local motion planner, the UAV successfully completes a 160-meter outdoor long-range navigation task. This work demonstrates how the world-understanding capabilities of foundation models can be transferred to embodied navigation, enabling robots to benefit from the strong generalization ability of existing image generation models.

Keywords

Cite

@article{arxiv.2605.07496,
  title  = {PathPainter: Transferring the Generalization Ability of Image Generation Models to Embodied Navigation},
  author = {Yijin Wang and Yuru Tian and Xijie Huang and Weiqi Gai and Mo Zhu and Xin Zhou and Yuze Wu and Fei Gao},
  journal= {arXiv preprint arXiv:2605.07496},
  year   = {2026}
}

Comments

Work in the progress. 11 pages, 7 figures

R2 v1 2026-07-01T12:57:22.375Z