Attend2Pack:基于注意力深度强化学习的装箱问题求解方法
机器学习
2021-08-03 v2 人工智能
摘要
本文试图从学习的视角解决装箱问题(BPP)。基于自注意力编码与深度强化学习算法,我们针对这一感兴趣的任务提出了一种新的端到端学习模型。通过分解组合动作空间,并利用一种称为优先过采样(prioritized oversampling)的新训练技术(一种加速同策略学习的通用方案),我们在一系列实验设置中取得了最先进的性能。此外,尽管所提方法 attend2pack 面向离线 BPP,我们将其简化至严格的在线 BPP 设置,在该设置下同样能够达到最先进的性能。通过一组消融研究以及与一系列已有工作的对比,我们希望为这一研究领域提供一个有效的基线方法。
引用
@article{arxiv.2107.04333,
title = {Attend2Pack: Bin Packing through Deep Reinforcement Learning with Attention},
author = {Jingwei Zhang and Bin Zi and Xiaoyu Ge},
journal= {arXiv preprint arXiv:2107.04333},
year = {2021}
}
备注
Reinforcement Learning for Real Life (RL4RealLife) Workshop in the 38th International Conference on Machine Learning, 2021