在连续任务中通过顾问迁移领域知识
人工智能
2021-02-17 v1
摘要
强化学习(RL)的近期进展已在许多模拟环境中超越人类水平表现。然而,现有强化学习技术无法将已已知的领域特定知识显式地纳入学习过程。因此,智能体必须通过试错方法独立探索并学习领域知识,这既耗费时间又消耗资源以作出有效响应。为此,我们改进了深度确定性策略梯度(DDPG)算法以引入一名顾问,从而以预学习策略或预定义关系的形式整合领域知识,以增强智能体的学习过程。我们在 OpenAi Gym 基准任务上的实验表明,通过顾问整合领域知识可加快学习速度并改进策略以趋向更优极值。
引用
@article{arxiv.2102.08029,
title = {Transferring Domain Knowledge with an Adviser in Continuous Tasks},
author = {Rukshan Wijesinghe and Kasun Vithanage and Dumindu Tissera and Alex Xavier and Subha Fernando and Jayathu Samarawickrama},
journal= {arXiv preprint arXiv:2102.08029},
year = {2021}
}
备注
Accepted by the 25th Pacific-Asia Conference on Knowledge Discovery and Data Mining (PAKDD-2021)