中文

从仿真到缩比城市:基于自动驾驶车辆的交通控制零样本策略迁移

系统与控制 2019-02-26 v2 人工智能 机器人学

摘要

我们利用深度强化学习,训练了用于引导车队驶入环岛的自动驾驶车辆控制策略。使用面向微仿真器的深度强化学习库 Flow,我们训练了两种策略:一种在状态与动作空间中注入噪声,另一种无任何注入噪声。在仿真中,自动驾驶车辆为两种策略均学习到一种涌现的调控行为,即减速以允许更平滑的合流。随后我们将该策略不经任何调参直接迁移至特拉华大学缩比智能城市(UDSSC)——一个 1:25 比例的网联与自动驾驶车辆测试床。我们刻画了两种策略在缩比城市中的表现。结果表明,无噪声策略最终发生碰撞且仅偶尔进行调控;而注入噪声的策略持续执行调控行为且保持无碰撞,这表明噪声有助于零样本策略迁移。此外,迁移后的注入噪声策略在 UDSSC 中使平均行程时间减少 5%,最大行程时间减少 22%。控制器视频见 https://sites.google.com/view/iccps-policy-transfer。

关键词

引用

@article{arxiv.1812.06120,
  title  = {Simulation to Scaled City: Zero-Shot Policy Transfer for Traffic Control via Autonomous Vehicles},
  author = {Kathy Jang and Eugene Vinitsky and Behdad Chalaki and Ben Remer and Logan Beaver and Andreas Malikopoulos and Alexandre Bayen},
  journal= {arXiv preprint arXiv:1812.06120},
  year   = {2019}
}

备注

To be published at the International Conference on Cyber Physical Systems (ICCPS) 2019. 10 pages, 9 figures