中文

通过受保护属性检测与看法分类进行偏见分析与缓解

计算与语言 2025-09-04 v3

摘要

大型语言模型(LLMs)从大规模预训练中获取通用语言知识。然而,预训练数据主要由网络爬取的文本组成,包含不良社会偏见,这些偏见可能被LLMs 延续或放大。本研究提出一种高效且有效的标注流程来调查预训练语料中的社会偏见。我们的流程包括受保护属性检测以识别多样化人口统计特征,随后进行看法分类以分析对每种属性的语言极性。通过我们的实验,我们展示了我们的偏见分析和缓解措施的效果,重点关注最具代表性的预训练语料Common Crawl。

关键词

引用

@article{arxiv.2504.14212,
  title  = {Bias Analysis and Mitigation through Protected Attribute Detection and Regard Classification},
  author = {Takuma Udagawa and Yang Zhao and Hiroshi Kanayama and Bishwaranjan Bhattacharjee},
  journal= {arXiv preprint arXiv:2504.14212},
  year   = {2025}
}

备注

Accepted to EMNLP 2025 (Findings)