English
Related papers

Related papers: Foundation Models and Fair Use

200 papers

Machine learning systems require representations of the real world for training and testing - they require data, and lots of it. Collecting data at scale has logistical and ethical challenges, and synthetic data promises a solution to these…

Computers and Society · Computer Science 2024-05-06 Cedric Deslandes Whitney , Justin Norman

The technology of formal software verification has made spectacular advances, but how much does it actually benefit the development of practical software? Considerable disagreement remains about the practicality of building systems with…

Software Engineering · Computer Science 2026-01-21 Li Huang , Sophie Ebersold , Alexander Kogtenkov , Bertrand Meyer , Yinling Liu

Through a systematization of generative AI (GenAI) stakeholder goals and expectations, this work seeks to uncover what value different stakeholders see in their contributions to the GenAI supply line. This valuation enables us to understand…

Artificial Intelligence · Computer Science 2024-08-02 Amruta Mahuli , Asia Biega

The ethical need to protect AI-generated content has been a significant concern in recent years. While existing watermarking strategies have demonstrated success in detecting synthetic content (detection), there has been limited exploration…

Cryptography and Security · Computer Science 2024-07-17 Rui Min , Sen Li , Hongyang Chen , Minhao Cheng

The release of ChatGPT, Gemini, and other large language model has drawn huge interests on foundations models. There is a broad consensus that foundations models will be the fundamental building blocks for future AI systems. However, there…

Computation and Language · Computer Science 2024-07-17 Qinghua Lu , Liming Zhu , Xiwei Xu , Zhenchang Xing , Jon Whittle

With the emerging application of Federated Learning (FL) in finance, hiring and healthcare, FL models are regulated to be fair, preventing disparities with respect to legally protected attributes such as race or gender. Two concepts of…

Machine Learning · Computer Science 2025-04-02 Yuying Duan , Gelei Xu , Yiyu Shi , Michael Lemmon

Generative AI models are capable of performing a wide variety of tasks that have traditionally required creativity and human understanding. During training, they learn patterns from existing data and can subsequently generate new content…

Pre-training, which utilizes extensive and varied datasets, is a critical factor in the success of Large Language Models (LLMs) across numerous applications. However, the detailed makeup of these datasets is often not disclosed, leading to…

Cryptography and Security · Computer Science 2024-01-02 Haodong Li , Gelei Deng , Yi Liu , Kailong Wang , Yuekang Li , Tianwei Zhang , Yang Liu , Guoai Xu , Guosheng Xu , Haoyu Wang

In a world of daily emerging scientific inquisition and discovery, the prolific launch of machine learning across industries comes to little surprise for those familiar with the potential of ML. Neither so should the congruent expansion of…

Artificial Intelligence · Computer Science 2021-12-13 Brianna Richardson , Juan E. Gilbert

Over the last years, topic modeling has emerged as a powerful technique for organizing and summarizing big collections of documents or searching for particular patterns in them. However, privacy concerns may arise when cross-analyzing data…

Machine Learning · Computer Science 2023-06-13 Lorena Calvo-Bartolomé , Jerónimo Arenas-García

Copyright infringement may occur when a generative model produces samples substantially similar to some copyrighted data that it had access to during the training phase. The notion of access usually refers to including copyrighted samples…

Machine Learning · Computer Science 2024-06-05 Yiwei Lu , Matthew Y. R. Yang , Zuoqiu Liu , Gautam Kamath , Yaoliang Yu

A growing ecosystem of large, open-source foundation models has reduced the labeled data and technical expertise necessary to apply machine learning to many new problems. Yet foundation models pose a clear dual-use risk, indiscriminately…

Machine Learning · Computer Science 2023-08-10 Peter Henderson , Eric Mitchell , Christopher D. Manning , Dan Jurafsky , Chelsea Finn

Are there any conditions under which a generative model's outputs are guaranteed not to infringe the copyrights of its training data? This is the question of "provable copyright protection" first posed by Vyas, Kakade, and Barak (ICML…

Cryptography and Security · Computer Science 2026-02-26 Aloni Cohen

Machine learning software is increasingly being used to make decisions that affect people's lives. But sometimes, the core part of this software (the learned model), behaves in a biased manner that gives undue advantages to a specific group…

Software Engineering · Computer Science 2020-10-07 Joymallya Chakraborty , Suvodeep Majumder , Zhe Yu , Tim Menzies

Deep learning (DL) models, especially those large-scale and high-performance ones, can be very costly to train, demanding a great amount of data and computational resources. Unauthorized reproduction of DL models can lead to copyright…

Cryptography and Security · Computer Science 2021-12-13 Jialuo Chen , Jingyi Wang , Tinglan Peng , Youcheng Sun , Peng Cheng , Shouling Ji , Xingjun Ma , Bo Li , Dawn Song

Obtaining a well-trained model involves expensive data collection and training procedures, therefore the model is a valuable intellectual property. Recent studies revealed that adversaries can `steal' deployed models even when they have no…

Cryptography and Security · Computer Science 2021-12-08 Yiming Li , Linghui Zhu , Xiaojun Jia , Yong Jiang , Shu-Tao Xia , Xiaochun Cao

Federated Foundation Models (FedFMs) represent a distributed learning paradigm that fuses general competences of foundation models as well as privacy-preserving capabilities of federated learning. This combination allows the large…

Federated learning has the potential to unlock siloed data and distributed resources by enabling collaborative model training without sharing private data. As more complex foundational models gain widespread use, the need to expand training…

Machine Learning · Computer Science 2025-09-08 Cosmin-Andrei Hatfaludi , Alex Serban

As deep image classification applications, e.g., face recognition, become increasingly prevalent in our daily lives, their fairness issues raise more and more concern. It is thus crucial to comprehensively test the fairness of these…

Machine Learning · Computer Science 2021-12-03 Peixin Zhang , Jingyi Wang , Jun Sun , Xinyu Wang

Practically all large language models have been pre-trained on data that is subject to global uncertainty related to copyright infringement and breach of contract. This creates potential risk for users and developers due to this uncertain…

Computation and Language · Computer Science 2025-04-11 Michael J Bommarito , Jillian Bommarito , Daniel Martin Katz