中文

BPL:偏置自适应偏好蒸馏学习用于推荐系统

机器学习 2025-10-21 v1 人工智能 信息检索

摘要

推荐系统受偏置影响,导致收集的反馈未能完整揭示用户偏好。虽然去偏置学习已广泛研究,但主要关注由随机曝光项目模拟的特定(称为反事实)测试环境,显著降低了基于实际用户-项目交互的典型(称为事实)测试环境中的准确性。事实上,不同的测试环境强调不同的优势:反事实测试强调用户长期满意度,而事实测试关注平台上后续用户行为的预测。因此,具有同时优于两种测试的模型更为理想。本文引入一种新型学习框架,称为偏置自适应偏好蒸馏学习(BPL),通过双重蒸馏策略逐步揭示用户偏好。这些蒸馈策略旨在在事实和反事实测试环境中实现高性能。采用一种特殊形式的教师-学生蒸馏,BPL 保留了与收集反馈一致的准确偏好知识,从而在事实测试中实现高性能。此外,通过可靠过滤的自蒸馏,BPL 在训练过程中不断细化其知识。这使得模型能在更广泛的用户-项目组合上提供更准确的预测,从而提高反事实测试中的性能。全面实验验证了 BPL 在事实和反事实测试中的有效性。我们的实现可访问:https://github.com/SeongKu-Kang/BPL。

关键词

引用

@article{arxiv.2510.16076,
  title  = {BPL: Bias-adaptive Preference Distillation Learning for Recommender System},
  author = {SeongKu Kang and Jianxun Lian and Dongha Lee and Wonbin Kweon and Sanghwan Jang and Jaehyun Lee and Jindong Wang and Xing Xie and Hwanjo Yu},
  journal= {arXiv preprint arXiv:2510.16076},
  year   = {2025}
}

备注

\c{opyright} 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works