English
Related papers

Related papers: The Data Provenance Initiative: A Large Scale Audi…

200 papers

Progress in AI is driven largely by the scale and quality of training data. Despite this, there is a deficit of empirical analysis examining the attributes of well-established datasets beyond text. In this work we conduct the largest and…

Many AI companies are training their large language models (LLMs) on data without the permission of the copyright owners. The permissibility of doing so varies by jurisdiction: in countries like the EU and Japan, this is allowed under…

Permissive licenses like MIT, Apache-2.0, and BSD-3-Clause dominate open-source AI, signaling that artifacts like models, datasets, and code can be freely used, modified, and redistributed. However, these licenses carry mandatory…

Machine Learning · Computer Science 2026-02-10 James Jewitt , Gopi Krishnan Rajbahadur , Hao Li , Bram Adams , Ahmed E. Hassan

This paper argues that a dataset's legal risk cannot be accurately assessed by its license terms alone; instead, tracking dataset redistribution and its full lifecycle is essential. However, this process is too complex for legal experts to…

Computers and Society · Computer Science 2025-03-17 Jaekyeom Kim , Sungryull Sohn , Gerrard Jeongwon Jo , Jihoon Choi , Kyunghoon Bae , Hwayoung Lee , Yongmin Park , Honglak Lee

Hidden license conflicts in the open-source AI ecosystem pose serious legal and ethical risks, exposing organizations to potential litigation and users to undisclosed risk. However, the field lacks a data-driven understanding of how…

Software Engineering · Computer Science 2025-09-15 James Jewitt , Hao Li , Bram Adams , Gopi Krishnan Rajbahadur , Ahmed E. Hassan

New capabilities in foundation models are owed in large part to massive, widely-sourced, and under-documented training data collections. Existing practices in data collection have led to challenges in tracing authenticity, verifying…

Artificial Intelligence · Computer Science 2024-09-04 Shayne Longpre , Robert Mahari , Naana Obeng-Marnu , William Brannon , Tobin South , Katy Gero , Sandy Pentland , Jad Kabbara

The rapid advancement of general-purpose AI models has increased concerns about copyright infringement in training data, yet current regulatory frameworks remain predominantly reactive rather than proactive. This paper examines the…

Computers and Society · Computer Science 2026-01-21 Mariia Kyrychenko , Mykyta Mudryi , Markiyan Chaklosh

Publicly available datasets are one of the key drivers for commercial AI software. The use of publicly available datasets is governed by dataset licenses. These dataset licenses outline the rights one is entitled to on a given dataset and…

Machine Learning · Computer Science 2022-04-12 Gopi Krishnan Rajbahadur , Erika Tuck , Li Zi , Dayi Lin , Boyuan Chen , Zhen Ming , Jiang , Daniel M. German

Does the training of large language models potentially infringe upon code licenses? Furthermore, are there any datasets available that can be safely used for training these models without violating such licenses? In our study, we assess the…

Software Engineering · Computer Science 2024-03-25 Jonathan Katzy , Răzvan-Mihai Popescu , Arie van Deursen , Maliheh Izadi

The rapid advancement of general-purpose AI models has increased concerns about copyright infringement in training data, yet current regulatory frameworks remain predominantly reactive rather than proactive. This paper examines the…

Computers and Society · Computer Science 2026-01-21 Mariia Kyrychenko , Mykyta Mudryi , Markiyan Chaklosh

The widespread use of Large Language Models (LLMs) raises critical concerns regarding the unauthorized inclusion of copyrighted content in training data. Existing detection frameworks, such as DE-COP, are computationally intensive, and…

Artificial Intelligence · Computer Science 2026-03-20 David Szczecina , Senan Gaffori , Edmond Li

Generative AI is becoming increasingly prevalent in creative fields, sparking urgent debates over how current copyright laws can keep pace with technological innovation. Recent controversies of AI models generating near-replicas of…

Machine Learning · Computer Science 2025-07-01 Archer Amon , Zhipeng Yin , Zichong Wang , Avash Palikhe , Wenbin Zhang

As datasets become critical assets in modern machine learning systems, ensuring robust copyright protection has emerged as an urgent challenge. Traditional legal mechanisms often fail to address the technical complexities of digital data…

Cryptography and Security · Computer Science 2025-09-09 Kun Li , Cheng Wang , Minghui Xu , Yue Zhang , Xiuzhen Cheng

We investigate the contents of web-scraped data for training AI systems, at sizes where human dataset curators and compilers no longer manually annotate every sample. Building off of prior privacy concerns in machine learning models, we…

Cryptography and Security · Computer Science 2026-04-08 Rachel Hong , Jevan Hutson , William Agnew , Imaad Huda , Tadayoshi Kohno , Jamie Morgenstern

Large Language Models (LLMs) have become central in academia and industry, raising concerns about privacy, transparency, and misuse. A key issue is the trustworthiness of proprietary models, with open-sourcing often proposed as a solution.…

Software Engineering · Computer Science 2025-01-29 Domen Vake , Bogdan Šinik , Jernej Vičič , Aleksandar Tošić

Speech AI Technologies are largely trained on publicly available datasets or by the massive web-crawling of speech. In both cases, data acquisition focuses on minimizing collection effort, without necessarily taking the data subjects'…

Computers and Society · Computer Science 2023-05-04 Orestis Papakyriakopoulos , Alice Xiang

Practically all large language models have been pre-trained on data that is subject to global uncertainty related to copyright infringement and breach of contract. This creates potential risk for users and developers due to this uncertain…

Computation and Language · Computer Science 2025-04-11 Michael J Bommarito , Jillian Bommarito , Daniel Martin Katz

In the rapidly evolving field of artificial intelligence (AI), mapping innovation patterns and understanding effective technology transfer from research to applications are essential for economic growth. However, existing data…

Databases · Computer Science 2025-06-02 Haixing Gong , Hui Zou , Xingzhou Liang , Shiyuan Meng , Pinlong Cai , Xingcheng Xu , Jingjing Qu

Artificial Intelligence (AI) has made its way into various scientific fields, providing astonishing improvements over existing algorithms for a wide variety of tasks. In recent years, there have been severe concerns over the trustworthiness…

Machine Learning · Computer Science 2024-08-20 Surbhi Mittal , Kartik Thakral , Richa Singh , Mayank Vatsa , Tamar Glaser , Cristian Canton Ferrer , Tal Hassner

The internet has become the main source of data to train modern text-to-image or vision-language models, yet it is increasingly unclear whether web-scale data collection practices for training AI systems adequately respect data owners'…

Computers and Society · Computer Science 2026-04-10 Chung Peng Lee , Rachel Hong , Harry H. Jiang , Aster Plotnik , William Agnew , Jamie Morgenstern
‹ Prev 1 2 3 10 Next ›