Deriving Rewards for Reinforcement Learning from Symbolic Behaviour Descriptions of Bipedal Walking
Abstract
Generating physical movement behaviours from their symbolic description is a long-standing challenge in artificial intelligence (AI) and robotics, requiring insights into numerical optimization methods as well as into formalizations from symbolic AI and reasoning. In this paper, a novel approach to finding a reward function from a symbolic description is proposed. The intended system behaviour is modelled as a hybrid automaton, which reduces the system state space to allow more efficient reinforcement learning. The approach is applied to bipedal walking, by modelling the walking robot as a hybrid automaton over state space orthants, and used with the compass walker to derive a reward that incentivizes following the hybrid automaton cycle. As a result, training times of reinforcement learning controllers are reduced while final walking speed is increased. The approach can serve as a blueprint how to generate reward functions from symbolic AI and reasoning.
Keywords
Cite
@article{arxiv.2312.10328,
title = {Deriving Rewards for Reinforcement Learning from Symbolic Behaviour Descriptions of Bipedal Walking},
author = {Daniel Harnack and Christoph Lüth and Lukas Gross and Shivesh Kumar and Frank Kirchner},
journal= {arXiv preprint arXiv:2312.10328},
year = {2023}
}
Comments
To appear in 62nd IEEE Conference on Decision and Control (CDC). For supplemental material, see here https://dfki-ric-underactuated-lab.github.io/orthant_rewards_biped_rl/