中文

基于目标条件强化学习的中段物流

机器学习 2026-05-05 v1 机器学习

摘要

中段物流描述了将包裹路由 through 由卡车连接、容量有限的枢纽网络的问题。我们将其重新表述为多目标目标条件马尔可夫决策过程 (MDP)。我们的方法结合了图神经网络和模型无关强化学习,从环境状态中提取小型特征图。

关键词

引用

@article{arxiv.2605.02461,
  title  = {Middle-mile logistics through the lens of goal-conditioned reinforcement learning},
  author = {Onno Eberhard and Thibaut Cuvelier and Michal Valko and Bruno De Backer},
  journal= {arXiv preprint arXiv:2605.02461},
  year   = {2026}
}

备注

Published at Neural Information Processing Systems (NeurIPS) 2023 Workshop on Goal-Conditioned Reinforcement Learning