中文

AGR:面向 LLMs 模仿偏置的年龄组公平性奖励

机器学习 2024-09-09 v1 人工智能 计算与语言

摘要

大语言模型可能存在年龄偏见,导致不同年龄组个体受到不平等的对待。虽然已有大量研究针对种族和性别偏见,但年龄偏见鲜有探讨。缺乏针对年龄偏见的指令调优和偏好数据集严重阻碍了其检测与衡量,而现有微调方法也很少关注年龄相关的公平性。本文构建了用于 RLHF 的年龄偏见偏好数据集和指令调优数据集。我们提出 ARG(Age Group fairness Reward),用于降低 LLMs 在不同年龄组间响应质量的差异。大量实验表明,该奖励显著提升了响应准确性,并减少了跨年龄组的性能差距。我们的源代码和数据集已提供链接。

关键词

引用

@article{arxiv.2409.04340,
  title  = {AGR: Age Group fairness Reward for Bias Mitigation in LLMs},
  author = {Shuirong Cao and Ruoxi Cheng and Zhiqiang Wang},
  journal= {arXiv preprint arXiv:2409.04340},
  year   = {2024}
}

备注

The first two authors contributed equally to this work. Corresponding to Zhiqiang Wang. ACKNOWLEDGMENT: we would like to thank the computing resources support from the State Key Laboratory of New Computer Software Technologies at Nanjing University