Testing on real machines is indispensable for robotic control algorithms. In the context of learning-based algorithms, especially VLA models, demand for large-scale evaluation, i.e. testing a large number of models on a large number of tasks, is becoming increasingly urgent. However, doing this right is highly non-trivial, especially when scalability and reproducibility is taken into account. In this report, we describe our methodology for constructing RoboChallenge, an online evaluation system to test robotic control algorithms, and our survey of recent state-of-the-art VLA models using our initial benchmark Table30.
@article{arxiv.2510.17950,
title = {RoboChallenge: Large-scale Real-robot Evaluation of Embodied Policies},
author = {Adina Yakefu and Bin Xie and Chongyang Xu and Enwen Zhang and Erjin Zhou and Fan Jia and Haitao Yang and Haoqiang Fan and Haowei Zhang and Hongyang Peng and Jing Tan and Junwen Huang and Kai Liu and Kaixin Liu and Kefan Gu and Qinglun Zhang and Ruitao Zhang and Saike Huang and Shen Cheng and Shuaicheng Liu and Tiancai Wang and Tiezhen Wang and Wei Sun and Wenbin Tang and Yajun Wei and Yang Chen and Youqiang Gui and Yucheng Zhao and Yunchao Ma and Yunfei Wei and Yunhuan Yang and Yutong Guo and Ze Chen and Zhengyuan Du and Ziheng Zhang and Ziming Liu and Ziwei Yan},
journal= {arXiv preprint arXiv:2510.17950},
year = {2025}
}
Comments
Authors are listed in alphabetical order. The official website is located at https://robochallenge.ai