用于不平衡分类与强化学习探索的 Scope 损失
机器学习
2023-08-09 v1 人工智能
摘要
我们证明了强化学习问题与监督分类问题之间的等价性。我们随之将强化学习中的探索-利用权衡等同于监督分类中的数据集不平衡问题,并发现二者在应对方式上的相似性。基于对上述问题的分析,我们推导出一种用于强化学习与监督分类的新型损失函数。Scope Loss 是我们的新损失函数,其调节梯度以防止由过度利用与数据集不平衡导致的性能损失,且无需任何调参。我们在一组基准强化学习任务与一个偏斜分类数据集上将 Scope Loss 与 SOTA 损失函数进行对比,表明 Scope Loss 优于其他损失函数。
引用
@article{arxiv.2308.04024,
title = {Scope Loss for Imbalanced Classification and RL Exploration},
author = {Hasham Burhani and Xiao Qi Shi and Jonathan Jaegerman and Daniel Balicki},
journal= {arXiv preprint arXiv:2308.04024},
year = {2023}
}
备注
11 pages, 2 figures, under review for NeurIPS 2023