MotionLM:将多智能体运动预测作为语言建模
计算机视觉与模式识别
2023-09-29 v1 人工智能
机器学习
机器人学
摘要
对道路智能体未来行为的可靠预测是自动驾驶车辆安全规划的关键组成部分。在此,我们将连续轨迹表示为离散运动 token 序列,并将多智能体运动预测转化为在该域上的语言建模任务。我们的模型 MotionLM 具有若干优势:首先,它不需要锚点或显式潜变量优化来学习多模态分布,而是利用单一标准语言建模目标,最大化序列 token 上的平均对数概率。其次,我们的方法绕开了事后交互启发式——即先生成个体智能体轨迹再进行交互打分,而是以单一自回归解码过程对交互式智能体未来产生联合分布。此外,该模型的序列分解支持时间因果条件滚动预测。所提方法在 Waymo Open Motion Dataset 上确立了多智能体运动预测的新 SOTA 性能,在交互式挑战排行榜上排名第 1。
引用
@article{arxiv.2309.16534,
title = {MotionLM: Multi-Agent Motion Forecasting as Language Modeling},
author = {Ari Seff and Brian Cera and Dian Chen and Mason Ng and Aurick Zhou and Nigamaa Nayakanti and Khaled S. Refaat and Rami Al-Rfou and Benjamin Sapp},
journal= {arXiv preprint arXiv:2309.16534},
year = {2023}
}
备注
To appear at the International Conference on Computer Vision (ICCV) 2023