中文

用于上下文多臂老虎机问题的神经网络委员会

神经与进化计算 2014-09-30 v1 机器学习

摘要

本文提出了一种新的上下文多臂老虎机(contextual bandit)算法 NeuralBandit,该算法不需要关于上下文和奖励平稳性的假设。我们训练多个神经网络来建模已知上下文下的奖励价值。基于多专家(multi-experts)方法,提出了两种变体以在线选择多层感知机的参数。所提出的算法在具有和不具有奖励平稳性的大型数据集上均通过了成功测试。

关键词

引用

@article{arxiv.1409.8191,
  title  = {A Neural Networks Committee for the Contextual Bandit Problem},
  author = {Robin Allesiardo and Raphael Feraud and Djallel Bouneffouf},
  journal= {arXiv preprint arXiv:1409.8191},
  year   = {2014}
}

备注

21st International Conference on Neural Information Processing