中文

InnateCoder:利用基础模型学习程序化选项

机器学习 2025-05-20 v1

摘要

除了迁移学习情境之外,强化学习智能体会从零开始启动其学习过程。因此,这些智能体必须经历缓慢的过程,才能学习甚至最基本的技能来解决问题。本文我们提出了 InnateCoder,一个利用 human knowledge 编码于基础模型中以提供以程序化策略形式编码的“先天技能”的系统,即时间扩展动作或 options。与现有学习 options 的方法不同,InnateCoder是在zeroshot情境下从generalhumanknowledge编码于基础模型中学习这些options,而不是从与环境交互中获得的knowledge。然后,InnateCoder 是在 zero-shot 情境下从 general human knowledge 编码于基础模型中学习这些 options,而不是从与环境交互中获得的 knowledge。然后,InnateCoder 通过将编码这些 options 的 programs 组合成更大更复杂的 programs 来 search 程序化策略。我们假设,InnateCoder的学习和使用options的方式可以提高当前方法forlearningprogrammaticpoliciessampleefficiency。在MicroRTSKareltheRobot上的empiricalresults支持了我们的假设,因为它们显示,InnateCoder 的学习和使用 options 的方式可以提高当前方法 for learning programmatic policies 的 sample efficiency。在 MicroRTS 和 Karel the Robot 上的 empirical results 支持了我们的假设,因为它们显示,InnateCoder 比不使用 options 或从经验中学习 options 的版本在 sample efficiency 上更佳。

关键词

引用

@article{arxiv.2505.12508,
  title  = {InnateCoder: Learning Programmatic Options with Foundation Models},
  author = {Rubens O. Moraes and Quazi Asif Sadmine and Hendrik Baier and Levi H. S. Lelis},
  journal= {arXiv preprint arXiv:2505.12508},
  year   = {2025}
}

备注

Accepted at IJCAI 2025