English

A Safety Modulator Actor-Critic Method in Model-Free Safe Reinforcement Learning and Application in UAV Hovering

Artificial Intelligence 2024-10-10 v1 Machine Learning Robotics

Abstract

This paper proposes a safety modulator actor-critic (SMAC) method to address safety constraint and overestimation mitigation in model-free safe reinforcement learning (RL). A safety modulator is developed to satisfy safety constraints by modulating actions, allowing the policy to ignore safety constraint and focus on maximizing reward. Additionally, a distributional critic with a theoretical update rule for SMAC is proposed to mitigate the overestimation of Q-values with safety constraints. Both simulation and real-world scenarios experiments on Unmanned Aerial Vehicles (UAVs) hovering confirm that the SMAC can effectively maintain safety constraints and outperform mainstream baseline algorithms.

Keywords

Cite

@article{arxiv.2410.06847,
  title  = {A Safety Modulator Actor-Critic Method in Model-Free Safe Reinforcement Learning and Application in UAV Hovering},
  author = {Qihan Qi and Xinsong Yang and Gang Xia and Daniel W. C. Ho and Pengyang Tang},
  journal= {arXiv preprint arXiv:2410.06847},
  year   = {2024}
}
R2 v1 2026-06-28T19:14:20.391Z