中文

用于自动驾驶按需出行系统的混合多智能体深度强化学习

机器学习 2023-05-11 v2 多智能体系统 系统与控制 系统与控制

摘要

我们考虑一个序贯决策问题:为一个利润最大的自动驾驶按需出行系统运营商做出主动的请求分配与拒绝决策。我们将此问题形式化为马尔可夫决策过程,并提出一种多智能体 Soft Actor-Critic 与加权二分匹配的 novel 组合,以获得具有前瞻性的控制策略。由此,我们对运营商原本难处理的动作空间进行因子分解,但仍获得全局协调的决策。基于真实世界出租车数据的实验表明,我们的方法在性能、稳定性和计算可处理性方面优于最先进的基准。

关键词

引用

@article{arxiv.2212.07313,
  title  = {Hybrid Multi-agent Deep Reinforcement Learning for Autonomous Mobility on Demand Systems},
  author = {Tobias Enders and James Harrison and Marco Pavone and Maximilian Schiffer},
  journal= {arXiv preprint arXiv:2212.07313},
  year   = {2023}
}

备注

20 pages, 7 figures, extended version of paper accepted at the 5th Learning for Dynamics & Control Conference (L4DC 2023)