中文

PokeAgent 挑战:大规模竞争与长程上下文学习

机器学习 2026-03-18 v2 人工智能

摘要

我们提出 PokeAgent 挑战,这是一个建立在 Pokemon 多智能体战斗系统和广泛的角色扮演游戏(RPG)环境上的大规模决策研究基准。部分可观测性、游戏论推理和长期计划 remain open problems for frontier AI,而几乎没有基准在现实条件下同时考验这三者。PokeAgent 通过两条互补轨道在规模上解决这些限制:我们的 Battling 轨道要求在部分可观测性下的竞争 Pokemon 战斗中进行战略推理和泛化;我们的 Speedrunning 轨道要求在 Pokemon RPG 环境中进行长期计划和序列决策。我们的 Battling 轨道提供 2000 万 以上战斗轨迹数据集,配套一套启发式、RL 和 LLM 基线,能够进行高水平竞争性游戏。我们的 Speedrunning 轨道提供首个 RPG speedrunning 的标准化评估框架,包括用于模块化、可重复比较基于 LLM 方法的开源多智能体编排系统。我们的 NeurIPS 2025 竞赛验证了资源质量和研究社区对 Pokemon 的兴趣,吸引了超过 100 支团队在两条轨道上竞争,获奖方案详述于本文。参赛提交和我们的基线结果显示,通用(LLM)、专家(RL)和 elite 人类表现之间存在显著差距。针对 BenchPress 评估矩阵的分析表明,Pokemon 战斗几乎与标准 LLM 基准正交,测度的是现有套件未捕捉的能力,使 Pokemon 成为可以推动 RL 和 LLM 研究前进的未解之谜。我们正在向一个活跃基准转变,为 Battling 提供实时排行榜,Speedrunning 提供自包含评估,网址为 https://pokeagentchallenge.com。

关键词

引用

@article{arxiv.2603.15563,
  title  = {The PokeAgent Challenge: Competitive and Long-Context Learning at Scale},
  author = {Seth Karten and Jake Grigsby and Tersoo Upaa and Junik Bae and Seonghun Hong and Hyunyoung Jeong and Jaeyoon Jung and Kun Kerdthaisong and Gyungbo Kim and Hyeokgi Kim and Yujin Kim and Eunju Kwon and Dongyu Liu and Patrick Mariglia and Sangyeon Park and Benedikt Schink and Xianwei Shi and Anthony Sistilli and Joseph Twin and Arian Urdu and Matin Urdu and Qiao Wang and Ling Wu and Wenli Zhang and Kunsheng Zhou and Stephanie Milani and Kiran Vodrahalli and Amy Zhang and Fei Fang and Yuke Zhu and Chi Jin},
  journal= {arXiv preprint arXiv:2603.15563},
  year   = {2026}
}

备注

41 pages, 26 figures, 5 tables. NeurIPS 2025 Competition Track