English

VistaBot: View-Robust Robot Manipulation via Spatiotemporal-Aware View Synthesis

Robotics 2026-04-24 v1

Abstract

Recently, end-to-end robotic manipulation models have gained significant attention for their generalizability and scalability. However, they often suffer from limited robustness to camera viewpoint changes when training with a fixed camera. In this paper, we propose VistaBot, a novel framework that integrates feed-forward geometric models with video diffusion models to achieve view-robust closed-loop manipulation without requiring camera calibration at test time. Our approach consists of three key components: 4D geometry estimation, view synthesis latent extraction, and latent action learning. VistaBot is integrated into both action-chunking (ACT) and diffusion-based (π0\pi_0) policies and evaluated across simulation and real-world tasks. We further introduce the View Generalization Score (VGS) as a new metric for comprehensive evaluation of cross-view generalization. Results show that VistaBot improves VGS by 2.79×\times and 2.63×\times over ACT and π0\pi_0, respectively, while also achieving high-quality novel view synthesis. Our contributions include a geometry-aware synthesis model, a latent action planner, a new benchmark metric, and extensive validation across diverse environments. The code and models will be made publicly available.

Keywords

Cite

@article{arxiv.2604.21914,
  title  = {VistaBot: View-Robust Robot Manipulation via Spatiotemporal-Aware View Synthesis},
  author = {Songen Gu and Yuhang Zheng and Weize Li and Yupeng Zheng and Yating Feng and Xiang Li and Yilun Chen and Pengfei Li and Wenchao Ding},
  journal= {arXiv preprint arXiv:2604.21914},
  year   = {2026}
}

Comments

This paper has been accepted to ICRA 2026

R2 v1 2026-07-01T12:32:52.036Z