中文

层次化通用价值函数逼近器

计算与语言 2024-10-14 v1

摘要

对于构建多目标强化学习价值函数的通用逼近器已取得重要进展——这们是估计状态参数化长期回报的关键元素。我们将其扩展到层次化强化学习,采用 options 框架,通过引入层次化通用价值函数逼近器(H-UVFAs)来实现。这使我们能够利用在时间抽象设置下预期的扩展性、规划和泛化优势。我们发展了监督学习和强化学习方法,用于学习状态、目标、选项和动作的嵌入向量,这两个层次化价值函数:Q(s,g,o;θ)Q(s, g, o; \theta)Q(s,g,o,a;θ)Q(s, g, o, a; \theta)。最后我们展示了 H-UVFAs 的泛化能力,并表明它们优于相应的 UVFAs。

关键词

引用

@article{arxiv.2410.08996,
  title  = {Hypothesis-only Biases in Large Language Model-Elicited Natural Language Inference},
  author = {Grace Proebsting and Adam Poliak},
  journal= {arXiv preprint arXiv:2410.08996},
  year   = {2024}
}