中文

威胁下的强化学习

机器学习 2019-10-28 v2 人工智能 密码学与安全 机器学习

摘要

在若干强化学习(RL)场景中,主要是在安全环境下,可能存在试图干扰奖励生成过程的对手。本文引入威胁马尔可夫决策过程(TMDPs),为 RL 中决策制定者抵御潜在对手提供支持框架。此外,我们提出了一种 level-kk 思维方案,由此产生一种处理 TMDPs 的新学习框架。在介绍我们的框架并推导理论结果后,通过大量实验给出了相关经验证据,表明在智能体学习过程中考虑对手的益处。

关键词

引用

@article{arxiv.1809.01560,
  title  = {Reinforcement Learning under Threats},
  author = {Victor Gallego and Roi Naveiro and David Rios Insua},
  journal= {arXiv preprint arXiv:1809.01560},
  year   = {2019}
}

备注

Extends the verson published at the Proceedings of the AAAI Conference on Artificial Intelligence 33, https://www.aaai.org/ojs/index.php/AAAI/article/view/5106