离策略分布Q($\lambda$):无需重要性采样的分布强化学习
机器学习
2024-02-09 v1
摘要
我们引入了离策略分布Q(),这是离策略分布评估算法家族的新成员。离策略分布Q()不应用重要性采样进行离策略学习,这引入了与符号测度的有趣交互。这种独特性质使分布Q()区别于其他现有替代方案,如分布Retrace。我们刻画了分布Q()的算法性质,并通过表格实验验证了理论见解。我们展示了分布Q()-C51,即Q()与C51智能体的结合,在深度强化学习基准上展现出有希望的结果。
引用
@article{arxiv.2402.05766,
title = {Off-policy Distributional Q($\lambda$): Distributional RL without Importance Sampling},
author = {Yunhao Tang and Mark Rowland and Rémi Munos and Bernardo Ávila Pires and Will Dabney},
journal= {arXiv preprint arXiv:2402.05766},
year = {2024}
}