用于学习非特指强化的两步算法
统计力学
2009-10-31 v3 无序系统与神经网络
摘要
我们研究了一个基于赫鲁布斯规则的简单学习模型,以应对“延迟”和非特指的强化。尽管信息反馈的性质不特指,但仍可观察到对渐近完美泛化的收敛,其收敛速率虽然取决于学习参数,但却呈非普适方式。渐近收敛的速度甚至可以与赫鲁布斯学习相同,亦可能更慢。此外,对于 certain 参数设置范围,系统是否能进入渐近完美泛化范数,取决于初始条件。
引用
@article{arxiv.cond-mat/9902354,
title = {A two step algorithm for learning from unspecific reinforcement},
author = {Reimer Kuehn and Ion-Olimpiu Stamatescu},
journal= {arXiv preprint arXiv:cond-mat/9902354},
year = {2009}
}
备注
13 pages LaTeX, 4 figures, note on biologically motivated stochastic variant of the algorithm added