中文

基于印度法律数据训练的模型是否公平?

计算与语言 2024-05-15 v3 计算机与社会

摘要

语言技术与人工智能近期的进展与应用已在法律、医疗与心理健康等多个领域取得诸多成功。基于AI的语言模型,如判决预测,近来已被提出用于法律行业。然而,这些模型深受从训练数据中习得的社交偏见所困。尽管偏见与公平性已在NLP领域得到研究,多数研究主要立足于西方语境。在本文中,我们呈现了从印度视角出发对法律领域公平性的初步考察。我们凸显了在 Hindi 法律文档上训练的模型于保释预测任务中习得算法偏见的传播。我们利用人口均等的差异(demographic parity)评估公平性差距,表明为保释预测任务训练的决策树模型在关联印度教徒与穆斯林的输入特征间存在0.237的总体公平性差异。此外,我们强调需要在印度语境下专门聚焦法律行业AI应用的公平性/偏见方向开展进一步研究。

关键词

引用

@article{arxiv.2303.07247,
  title  = {Are Models Trained on Indian Legal Data Fair?},
  author = {Sahil Girhepuje and Anmol Goel and Gokul S Krishnan and Shreya Goyal and Satyendra Pandey and Ponnurangam Kumaraguru and Balaraman Ravindran},
  journal= {arXiv preprint arXiv:2303.07247},
  year   = {2024}
}

备注

Presented at the Symposium on AI and Law (SAIL) 2023