English

Learning Robust Policies via Interpretable Hamilton-Jacobi Reachability-Guided Disturbances

Robotics 2024-10-01 v1

Abstract

Deep Reinforcement Learning (RL) has shown remarkable success in robotics with complex and heterogeneous dynamics. However, its vulnerability to unknown disturbances and adversarial attacks remains a significant challenge. In this paper, we propose a robust policy training framework that integrates model-based control principles with adversarial RL training to improve robustness without the need for external black-box adversaries. Our approach introduces a novel Hamilton-Jacobi reachability-guided disturbance for adversarial RL training, where we use interpretable worst-case or near-worst-case disturbances as adversaries against the robust policy. We evaluated its effectiveness across three distinct tasks: a reach-avoid game in both simulation and real-world settings, and a highly dynamic quadrotor stabilization task in simulation. We validate that our learned critic network is consistent with the ground-truth HJ value function, while the policy network shows comparable performance with other learning-based methods.

Keywords

Cite

@article{arxiv.2409.19746,
  title  = {Learning Robust Policies via Interpretable Hamilton-Jacobi Reachability-Guided Disturbances},
  author = {Hanyang Hu and Xilun Zhang and Xubo Lyu and Mo Chen},
  journal= {arXiv preprint arXiv:2409.19746},
  year   = {2024}
}
R2 v1 2026-06-28T19:01:11.758Z