中文

基于 Lipschitz 多臂老虎机的 POMDP 高效采样用于连续空间运动规划

机器人学 2021-06-09 v1 机器学习

摘要

不确定性下的决策可建模为部分可观测马尔可夫决策过程 (POMDP)。求 POMDP 精确解通常在计算上不可行,但可通过基于采样的方法近似求解。这些基于采样的 POMDP 求解器依赖多臂老虎机 (MAB) 启发式,其假设不同动作的结果不相关。在某些应用中,如连续空间中的运动规划,相似动作产生相似结果。本文中,我们利用做出动作结果 Lipschitz 连续性假设的 MAB 启发式变体,以提升基于采样的规划方法效率。我们在自动驾驶运动规划背景下展示了该方法的有效性。

关键词

引用

@article{arxiv.2106.04206,
  title  = {Efficient Sampling in POMDPs with Lipschitz Bandits for Motion Planning in Continuous Spaces},
  author = {Ömer Şahin Taş and Felix Hauser and Martin Lauer},
  journal= {arXiv preprint arXiv:2106.04206},
  year   = {2021}
}

备注

In Proceedings of the IEEE Intelligent Vehicle Symposium 2021