中文

文本究竟归谁所有?探索 BigCode、知识产权与伦理

计算机与社会 2023-04-07 v1 人工智能

摘要

智能或生成式写作工具依赖于能够识别、总结、翻译和预测内容的大语言模型(LLM)。本立场论文探究用于训练大语言模型(LLM)的开放数据集的版权利益。我们的论文追问:在开放数据集上训练的 LLM 如何规避所用数据的版权利益?我们首先界定软件版权并追溯其历史。我们以 GitHub Copilot 作为挑战软件现代化的案例研究。我们的结论概述了生成式写作助手给版权带来的障碍,并为开发者、软件法律专家和普通用户在智能 LLM 驱动写作工具的背景下提供了一份实用的版权分析路线图。

关键词

引用

@article{arxiv.2304.02839,
  title  = {Whose Text Is It Anyway? Exploring BigCode, Intellectual Property, and Ethics},
  author = {Madiha Zahrah Choksi and David Goedicke},
  journal= {arXiv preprint arXiv:2304.02839},
  year   = {2023}
}

备注

3 pages, submitted to the Second Workshop on Intelligent and Interactive Writing Assistants co-located with the ACM CHI Conference on Human Factors in Computing Systems (CHI 2023)