中文

我是否应该公开我的数据集?可复现性与个人数据权利之间的权衡

计算机与社会 2022-11-02 v1

摘要

自然语言处理技术已帮助领域专家解决法律问题。法院文件的数字化可用性增加了研究人员的可能性,他们可以将其作为构建数据集的来源——而公开这些数据集符合计算研究中良好的可复现性实践。巴西等大型数字化法院系统很容易在此意义上被探索。然而,个人数据保护法律对数据暴露施加了限制,并规定了研究人员应当注意的原则。在涉及侵犯人权(例如我们作为关注示例详细阐述的性别歧视)的案件中必须特别谨慎。我们就该问题提出了法律与伦理考量,并为处理此类数据并决定是否公开的研究人员提供了指导方针。

关键词

引用

@article{arxiv.2211.00498,
  title  = {Should I disclose my dataset? Caveats between reproducibility and individual data rights},
  author = {Raysa M. Benatti and Camila M. L. Villarroel and Sandra Avila and Esther L. Colombini and Fabiana C. Severi},
  journal= {arXiv preprint arXiv:2211.00498},
  year   = {2022}
}

备注

10 pages, 2 figures. To be published in the 4th Workshop on Natural Legal Language Processing (NLLP 2022), co-located with the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP 2022)