English

Enhancing RL Safety with Counterfactual LLM Reasoning

Machine Learning 2024-09-17 v1

Abstract

Reinforcement learning (RL) policies may exhibit unsafe behavior and are hard to explain. We use counterfactual large language model reasoning to enhance RL policy safety post-training. We show that our approach improves and helps to explain the RL policy safety.

Keywords

Cite

@article{arxiv.2409.10188,
  title  = {Enhancing RL Safety with Counterfactual LLM Reasoning},
  author = {Dennis Gross and Helge Spieker},
  journal= {arXiv preprint arXiv:2409.10188},
  year   = {2024}
}