中文

情境赌博机中带情景式奖励的在线半监督学习

机器学习 2020-10-27 v2 人工智能 机器学习

摘要

我们考虑了一个由若干真实应用驱动的、具有情景式揭示奖励的在线学习新实际问题,其中情境在不同情景间非平稳,且奖励反馈并非总对决策智能体可用。针对此在线半监督学习设定,我们引入了背景情景式奖励LinUCB(BerlinUCB),该方案轻松将聚类作为自监督模块,以在未观测到奖励时提供有用的侧信息。我们在六种不同场景的平稳与非平稳环境中、多种数据集上的实验表明,所提方法相较标准情境赌博机具有明显优势。最后,我们介绍了一个该问题设定尤为有用的相关真实案例。

关键词

引用

@article{arxiv.2009.08457,
  title  = {Online Semi-Supervised Learning in Contextual Bandits with Episodic Reward},
  author = {Baihan Lin},
  journal= {arXiv preprint arXiv:2009.08457},
  year   = {2020}
}

备注

Proceeding of AJCAI 2020. This article supersedes our work arXiv:1802.00981 on contextual bandits in nonstationary setting, introduces a new problem setting with episodically revealed reward, and provides a novel solution by propagating pseudo-feedbacks to un-rewarded cases from self-supervision. Also check out our speaker diarization application of this algorithm at arXiv:2006.04376