English

REALM: A Real-to-Sim Validated Benchmark for Generalization in Robotic Manipulation

Robotics 2025-12-23 v1 Artificial Intelligence

Abstract

Vision-Language-Action (VLA) models empower robots to understand and execute tasks described by natural language instructions. However, a key challenge lies in their ability to generalize beyond the specific environments and conditions they were trained on, which is presently difficult and expensive to evaluate in the real-world. To address this gap, we present REALM, a new simulation environment and benchmark designed to evaluate the generalization capabilities of VLA models, with a specific emphasis on establishing a strong correlation between simulated and real-world performance through high-fidelity visuals and aligned robot control. Our environment offers a suite of 15 perturbation factors, 7 manipulation skills, and more than 3,500 objects. Finally, we establish two task sets that form our benchmark and evaluate the \pi_{0}, \pi_{0}-FAST, and GR00T N1.5 VLA models, showing that generalization and robustness remain an open challenge. More broadly, we also show that simulation gives us a valuable proxy for the real-world and allows us to systematically probe for and quantify the weaknesses and failure modes of VLAs. Project page: https://martin-sedlacek.com/realm

Keywords

Cite

@article{arxiv.2512.19562,
  title  = {REALM: A Real-to-Sim Validated Benchmark for Generalization in Robotic Manipulation},
  author = {Martin Sedlacek and Pavlo Yefanov and Georgy Ponimatkin and Jai Bardhan and Simon Pilc and Mederic Fourmy and Evangelos Kazakos and Cees G. M. Snoek and Josef Sivic and Vladimir Petrik},
  journal= {arXiv preprint arXiv:2512.19562},
  year   = {2025}
}

Comments

9 pages, 10 figures

R2 v1 2026-07-01T08:37:12.824Z