中文

程序性政策的来源:程序性空间与潜在空间

机器学习 2024-10-17 v1 人工智能

摘要

最近的研究引入了 LEAPS 和 HPRL 系统,这些系统学习 domain-specific languages 的潜在空间,用于为部分可观测马尔可夫决策过程 (POMDP) 定义 programmatic policies。这些系统在优化诸如 behavior loss 等损失函数的同时诱导潜在空间,这些损失函数旨在实现 program behavior 的 locality,即潜在空间中靠近的向量应对应于 similarly 行为的 programs。在本文中,我们表明,由 domain-specific language 诱导且无需训练的 programmatic space,其 behavior loss 的值与之前工作中呈现的潜在空间相似。此外,在 programmatic space 中搜索的算法在 LEAPS 和 HPRL 中表现显著优于。为了解释我们的结果,我们测量了两种空间对 local search 算法的 "friendliness"。我们发现,在潜在空间中搜索时,算法更容易停留在 local maxima 附近。这表明,programmatic space 的优化拓扑(由 reward function 与 neighborhood function 共同诱导)比潜在空间更有利于 search。这一结果为 programmatic space 的优越性能提供了解释。

关键词

引用

@article{arxiv.2410.12166,
  title  = {Reclaiming the Source of Programmatic Policies: Programmatic versus Latent Spaces},
  author = {Tales H. Carvalho and Kenneth Tjhia and Levi H. S. Lelis},
  journal= {arXiv preprint arXiv:2410.12166},
  year   = {2024}
}

备注

Published as a conference paper at ICLR 2024