中文

推荐中潜动作空间的探索与正则化

信息检索 2023-02-09 v2

摘要

在推荐系统中,强化学习方案因能捕捉长期用户-系统交互而有效提升了推荐性能。然而,推荐策略的动作空间是一个物品列表,在动态候选物品池下可能极其庞大。为克服该挑战,我们提出一种超 actor 与 critic 学习框架,其中策略将物品列表生成过程分解为超动作推断步骤与效果动作选择步骤。第一步将给定状态空间映射至向量化超动作空间,第二步基于超动作选择物品列表。为调节两个动作空间间的差异,我们设计了一个对齐模块及物品核映射函数以确保推断精度,并引入监督模块以稳定学习过程。我们在公开数据集上构建模拟环境,并实证表明我们的框架在推荐上优于标准 RL 基线。

关键词

引用

@article{arxiv.2302.03431,
  title  = {Exploration and Regularization of the Latent Action Space in Recommendation},
  author = {Shuchang Liu and Qingpeng Cai and Bowen Sun and Yuhao Wang and Ji Jiang and Dong Zheng and Kun Gai and Peng Jiang and Xiangyu Zhao and Yongfeng Zhang},
  journal= {arXiv preprint arXiv:2302.03431},
  year   = {2023}
}

备注

Proceedings of the ACM Web Conference 2023 (WWW '23), May 1--5, 2023, Austin, TX, USA