English

Galaxea Open-World Dataset and G0 Dual-System VLA Model

Robotics 2025-09-03 v1 Computer Vision and Pattern Recognition

Abstract

We present Galaxea Open-World Dataset, a large-scale, diverse collection of robot behaviors recorded in authentic human living and working environments. All demonstrations are gathered using a consistent robotic embodiment, paired with precise subtask-level language annotations to facilitate both training and evaluation. Building on this dataset, we introduce G0, a dual-system framework that couples a Vision-Language Model (VLM) for multimodal planning with a Vision-Language-Action (VLA) model for fine-grained execution. G0 is trained using a three-stage curriculum: cross-embodiment pre-training, single-embodiment pre-training, and task-specific post-training. A comprehensive benchmark spanning tabletop manipulation, few-shot learning, and long-horizon mobile manipulation, demonstrates the effectiveness of our approach. In particular, we find that the single-embodiment pre-training stage, together with the Galaxea Open-World Dataset, plays a critical role in achieving strong performance.

Keywords

Cite

@article{arxiv.2509.00576,
  title  = {Galaxea Open-World Dataset and G0 Dual-System VLA Model},
  author = {Tao Jiang and Tianyuan Yuan and Yicheng Liu and Chenhao Lu and Jianning Cui and Xiao Liu and Shuiqi Cheng and Jiyang Gao and Huazhe Xu and Hang Zhao},
  journal= {arXiv preprint arXiv:2509.00576},
  year   = {2025}
}

Comments

https://opengalaxea.github.io/G0/

R2 v1 2026-07-01T05:13:38.483Z