English

Reliable and Scalable Robot Policy Evaluation with Imperfect Simulators

Robotics 2025-10-07 v1 Artificial Intelligence Systems and Control Systems and Control

Abstract

Rapid progress in imitation learning, foundation models, and large-scale datasets has led to robot manipulation policies that generalize to a wide-range of tasks and environments. However, rigorous evaluation of these policies remains a challenge. Typically in practice, robot policies are often evaluated on a small number of hardware trials without any statistical assurances. We present SureSim, a framework to augment large-scale simulation with relatively small-scale real-world testing to provide reliable inferences on the real-world performance of a policy. Our key idea is to formalize the problem of combining real and simulation evaluations as a prediction-powered inference problem, in which a small number of paired real and simulation evaluations are used to rectify bias in large-scale simulation. We then leverage non-asymptotic mean estimation algorithms to provide confidence intervals on mean policy performance. Using physics-based simulation, we evaluate both diffusion policy and multi-task fine-tuned π0\pi_0 on a joint distribution of objects and initial conditions, and find that our approach saves over 2025%20-25\% of hardware evaluation effort to achieve similar bounds on policy performance.

Keywords

Cite

@article{arxiv.2510.04354,
  title  = {Reliable and Scalable Robot Policy Evaluation with Imperfect Simulators},
  author = {Apurva Badithela and David Snyder and Lihan Zha and Joseph Mikhail and Matthew O'Kelly and Anushri Dixit and Anirudha Majumdar},
  journal= {arXiv preprint arXiv:2510.04354},
  year   = {2025}
}
R2 v1 2026-07-01T06:18:14.328Z