推荐系统离线评估中的普遍缺陷
信息检索
2023-07-28 v1
摘要
尽管离线评估只是线上性能的不完美代理——由于推荐系统的交互性质——它在可预见的未来可能仍将是推荐系统研究中的主要评估方式,因为生产级推荐系统的专有性质阻碍了对于 A/B 测试设置的独立验证和线上结果的核实。因此,离线评估设置必须尽可能真实且无缺陷。遗憾的是,由于后续工作不加质疑地复制前人存在缺陷的评估设置,评估缺陷在当今推荐系统研究中相当常见。为提升推荐系统离线评估的质量,我们讨论了其中四种普遍缺陷以及研究者为何应当避免它们。
引用
@article{arxiv.2307.14951,
title = {Widespread Flaws in Offline Evaluation of Recommender Systems},
author = {Balázs Hidasi and Ádám Tibor Czapp},
journal= {arXiv preprint arXiv:2307.14951},
year = {2023}
}
备注
Appearing in the Proceedings of the 17th ACM Conference on Recommender Systems