English

Specifying Non-Markovian Rewards in MDPs Using LDL on Finite Traces (Preliminary Version)

Artificial Intelligence 2017-06-27 v1

Abstract

In Markov Decision Processes (MDPs), the reward obtained in a state depends on the properties of the last state and action. This state dependency makes it difficult to reward more interesting long-term behaviors, such as always closing a door after it has been opened, or providing coffee only following a request. Extending MDPs to handle such non-Markovian reward function was the subject of two previous lines of work, both using variants of LTL to specify the reward function and then compiling the new model back into a Markovian model. Building upon recent progress in the theories of temporal logics over finite traces, we adopt LDLf for specifying non-Markovian rewards and provide an elegant automata construction for building a Markovian model, which extends that of previous work and offers strong minimality and compositionality guarantees.

Keywords

Cite

@article{arxiv.1706.08100,
  title  = {Specifying Non-Markovian Rewards in MDPs Using LDL on Finite Traces (Preliminary Version)},
  author = {Ronen Brafman and Giuseppe De Giacomo and Fabio Patrizi},
  journal= {arXiv preprint arXiv:1706.08100},
  year   = {2017}
}