中文

MASS-DPO:基于Plackett--Luce模型的多负样本主动样本选择直接政策优化

机器学习 2026-05-12 v1

摘要

多负样本偏好优化延伸了直接偏好优化(DPO),通过在一个首选响应和多个被拒绝响应之间建立比较信号。然而,针对大规模负样本池的优化代价高昂,且许多候选因其对策略更新相似的影响而贡献冗余梯度。我们引入MASS-DPO,一种多负样本主动样本选择方法,推导针对Plackett--Luce模型的PL特有Fisher信息目标,用于在每个提示内选择紧凑、信息丰富的负样本子集。 resulting log-determinant objective selects negatives that contribute complementary information for policy updates, yielding compact subsets that retain the full pool's information while reducing redundancy. In practice, this favors negatives whose gradients cover different update directions, reducing redundant signal from near-duplicate candidates while preserving the most useful training information. Across four benchmarks spanning recommendation and multiple-choice QA and three model families, MASS-DPO consistently exceeds or matches existing methods in accuracy, improves Recall/NDCG and margin-based optimization dynamics, and delivers stronger alignment with substantially fewer negatives.

关键词

引用

@article{arxiv.2605.10784,
  title  = {MASS-DPO: Multi-negative Active Sample Selection for Direct Policy Optimization},
  author = {Rohan Surana and Xintong Li and Sheldon Yu and Yiran Jenny Shen and Chuhan Wang and Tong Yu and Prithviraj Ammanabrolu and Jingbo Shang and Julian McAuley and Junda Wu},
  journal= {arXiv preprint arXiv:2605.10784},
  year   = {2026}
}