中文

VI 与 VDRL 反馈函数的建模:通过计算模拟探寻规范法则

人工智能 2022-07-08 v2 定量方法

摘要

本文介绍一个名为 Beak 的 R 脚本,用于模拟行为与强化程序交互的速率。利用 Beak,我们模拟了可评估不同强化反馈函数(RFF)的数据。由于模拟提供海量数据样本,且更重要的是模拟行为不受其产生的强化所改变,因此得以前所未有的精度实现,并可系统性地改变行为。我们以意义、精度、简洁性和普适性为准则,比较了用于 RI 程序的不同 RFF。结果表明,RI 程序的最佳反馈函数由 Baum(1981)发表。我们还提出 Killeen(1975)所用模型可作为 RDRL 程序的可行反馈函数。我们认为 Beak 为深入理解强化程序铺平了道路,解决了有关程序定量特征尚待解决的问题,并可指导未来将程序作为理论与方法学工具使用的实验。

关键词

引用

@article{arxiv.2111.13943,
  title  = {Modeling VI and VDRL feedback functions: searching normative rules through computational simulation},
  author = {Paulo Sergio Panse Silveira and Jose de Oliveira Siqueira and Joao Lucas Bernardy and Jessica Santiago and Thiago Cersosimo Meneses and Bianca Sanches Portela and Marcelo Frota Benvenuti},
  journal= {arXiv preprint arXiv:2111.13943},
  year   = {2022}
}

备注

This is a revised version of the manuscript submitted for consideration for publication in the Journal of the Experimental Analysis of Behavior (JEAB) in July 6th, 2022. Supplemental material is available under SourceForge at https://sourceforge.net/projects/simpleschedules/