可提示游戏模型:基于掩码扩散模型的文本引导游戏模拟
摘要
神经视频游戏模拟器已成为生成和编辑视频的强大工具。其思想是将游戏表示为由智能体动作驱动的环境状态演化。虽然这种范式使用户能够逐动作地玩游戏,但其刚性排除了更具语义形式的控制。为克服此限制,我们用指定为一组自然语言动作和期望状态的提示来增强游戏模型。其结果——一个可提示游戏模型(PGM)——使用户能够通过以高层和低层动作序列提示来玩游戏。最引人入胜的是,我们的 PGM 解锁了导演模式,其中游戏通过以提示形式指定智能体的目标来游玩。这需要学习由我们的动画模型封装的“游戏 AI”,以使用高层约束导航场景、对抗对手并设计赢得一分的策略。为渲染所得状态,我们使用由我们的合成模型封装的组合式 NeRF 表征。为促进未来研究,我们呈现了新收集、标注和校准的 Tennis 与 Minecraft 数据集。我们的方法在渲染质量上显著优于现有神经视频游戏模拟器,并解锁了超出当前最先进水平能力范围的应用。我们的框架、数据和模型可在 https://snap-research.github.io/promptable-game-models/ 获取。
引用
@article{arxiv.2303.13472,
title = {Promptable Game Models: Text-Guided Game Simulation via Masked Diffusion Models},
author = {Willi Menapace and Aliaksandr Siarohin and Stéphane Lathuilière and Panos Achlioptas and Vladislav Golyanik and Sergey Tulyakov and Elisa Ricci},
journal= {arXiv preprint arXiv:2303.13472},
year = {2024}
}
备注
ACM Transactions on Graphics \c{opyright} Copyright is held by the owner/author(s) 2023. This is the author's version of the work. It is posted here for your personal use. Not for redistribution. The definitive Version of Record was published in ACM Transactions on Graphics, http://dx.doi.org/10.1145/3635705