中文

人类行为的 Bandit 模型:精神障碍中的奖赏处理

人工智能 2017-06-12 v1

摘要

受人类决策行为研究的启发,我们在此提出了一个用于多臂赌博机问题的通用参数化框架,该框架扩展了标准的 Thompson Sampling 方法,以纳入与多种神经和精神疾病相关的奖赏处理偏差,包括帕金森病、阿尔茨海默病、注意缺陷/多动障碍(ADHD)、成瘾和慢性疼痛。我们通过实验证明,所提出的参数化方法在多种数据集上通常能优于基线 Thompson Sampling。此外,从行为建模的角度来看,我们的参数化框架可被视为迈向统一计算模型的第一步,该模型旨在捕捉多种精神疾病中的奖赏处理异常。

关键词

引用

@article{arxiv.1706.02897,
  title  = {Bandit Models of Human Behavior: Reward Processing in Mental Disorders},
  author = {Djallel Bouneffouf and Irina Rish and Guillermo A. Cecchi},
  journal= {arXiv preprint arXiv:1706.02897},
  year   = {2017}
}

备注

Conference on Artificial General Intelligence, AGI-17