English
Related papers

Related papers: Mitigating Dataset Harms Requires Stewardship: Les…

200 papers

Despite extensive efforts to create fairer machine learning (ML) datasets, there remains a limited understanding of the practical aspects of dataset curation. Drawing from interviews with 30 ML dataset curators, we present a comprehensive…

Large organizations such as social media companies continually release data, for example user images. At the same time, these organizations leverage their massive corpora of released data to train proprietary models that give them an edge…

Cryptography and Security · Computer Science 2021-03-08 Liam Fowl , Ping-yeh Chiang , Micah Goldblum , Jonas Geiping , Arpit Bansal , Wojtek Czaja , Tom Goldstein

Dataset licensing is currently an issue in the development of machine learning systems. And in the development of machine learning systems, the most widely used are publicly available datasets. However, since the images in the publicly…

Software Engineering · Computer Science 2023-03-27 Junyu Chen , Norihiro Yoshida , Hiroaki Takada

The last two years have seen a rapid growth in concerns around the safety of large language models (LLMs). Researchers and practitioners have met these concerns by creating an abundance of datasets for evaluating and improving LLM safety.…

Computation and Language · Computer Science 2025-01-13 Paul Röttger , Fabio Pernisi , Bertie Vidgen , Dirk Hovy

The use of dialogue systems as a medium for human-machine interaction is an increasingly prevalent paradigm. A growing number of dialogue systems use conversation strategies that are learned from large datasets. There are well documented…

Computation and Language · Computer Science 2017-11-27 Peter Henderson , Koustuv Sinha , Nicolas Angelard-Gontier , Nan Rosemary Ke , Genevieve Fried , Ryan Lowe , Joelle Pineau

Data filtering strategies are a crucial component to develop safe Large Language Models (LLM), since they support the removal of harmful contents from pretraining datasets. There is a lack of research on the actual impact of these…

Computation and Language · Computer Science 2026-03-24 Marco Antonio Stranisci , Christian Hardmeier

In recent years, the rapid development of artificial intelligence (AI) systems has raised concerns about our ability to ensure their fairness, that is, how to avoid discrimination based on protected characteristics such as gender, race, or…

Computer Vision and Pattern Recognition · Computer Science 2025-03-06 Iris Dominguez-Catena , Daniel Paternain , Mikel Galar , MaryBeth Defrance , Maarten Buyl , Tijl De Bie

Machine learning datasets are powerful but unwieldy. Despite the fact that large datasets commonly contain problematic material--whether from a technical, legal, or ethical perspective--datasets are valuable resources when handled carefully…

Computers and Society · Computer Science 2025-01-28 Sarah Ciston , Mike Ananny , Kate Crawford

Speech AI Technologies are largely trained on publicly available datasets or by the massive web-crawling of speech. In both cases, data acquisition focuses on minimizing collection effort, without necessarily taking the data subjects'…

Computers and Society · Computer Science 2023-05-04 Orestis Papakyriakopoulos , Alice Xiang

LLMs produce harmful and undesirable behavior when trained on datasets containing even a small fraction of poisoned data. We demonstrate that GPT models remain vulnerable to fine-tuning on poisoned data, even when safeguarded by moderation…

Cryptography and Security · Computer Science 2025-07-18 Dillon Bowen , Brendan Murphy , Will Cai , David Khachaturov , Adam Gleave , Kellin Pelrine

Face recognition has achieved outstanding performance in the last decade with the development of deep learning techniques. Nowadays, the challenges in face recognition are related to specific scenarios, for instance, the performance under…

Computer Vision and Pattern Recognition · Computer Science 2023-01-18 Iurii Medvedev , Farhad Shadmand , Nuno Gonçalves

Large Language Models (LLMs) are foundational to AI advancements, facilitating applications like predictive text generation. Nonetheless, they pose risks by potentially memorizing and disseminating sensitive, biased, or copyrighted…

Artificial Intelligence · Computer Science 2024-03-26 Youyang Qu , Ming Ding , Nan Sun , Kanchana Thilakarathna , Tianqing Zhu , Dusit Niyato

Do large datasets provide value to psychologists? Without a systematic methodology for working with such datasets, there is a valid concern that analyses will produce noise artifacts rather than true effects. In this paper, we offer a way…

Computers and Society · Computer Science 2020-01-10 Mayank Agrawal , Joshua C. Peterson , Thomas L. Griffiths

As Large Language Models (LLMs) become integral to scientific workflows, concerns over the confidentiality and ethical handling of confidential data have emerged. This paper explores data exposure risks through LLM-powered scientific tools,…

Human-Computer Interaction · Computer Science 2025-04-15 Yashothara Shanmugarasa , Shidong Pan , Ming Ding , Dehai Zhao , Thierry Rakotoarivelo

Many AI companies are training their large language models (LLMs) on data without the permission of the copyright owners. The permissibility of doing so varies by jurisdiction: in countries like the EU and Japan, this is allowed under…

Advances in machine learning are closely tied to the creation of datasets. While data documentation is widely recognized as essential to the reliability, reproducibility, and transparency of ML, we lack a systematic empirical understanding…

Machine Learning · Computer Science 2024-01-26 Xinyu Yang , Weixin Liang , James Zou

Although essential to revealing biased performance, well intentioned attempts at algorithmic auditing can have effects that may harm the very populations these measures are meant to protect. This concern is even more salient while auditing…

Computers and Society · Computer Science 2020-01-07 Inioluwa Deborah Raji , Timnit Gebru , Margaret Mitchell , Joy Buolamwini , Joonseok Lee , Emily Denton

In an effort to mitigate the harms of large language models (LLMs), learning from human feedback (LHF) has been used to steer LLMs towards outputs that are intended to be both less harmful and more helpful. Despite the widespread adoption…

Computation and Language · Computer Science 2025-06-05 Khaoula Chehbouni , Jonathan Colaço Carr , Yash More , Jackie CK Cheung , Golnoosh Farnadi

In this paper we investigate problematic practices and consequences of large scale vision datasets. We examine broad issues such as the question of consent and justice as well as specific concerns such as the inclusion of verifiably…

Computers and Society · Computer Science 2020-07-27 Vinay Uday Prabhu , Abeba Birhane

Several high-profile events, such as the mass testing of emotion recognition systems on vulnerable sub-populations and using question answering systems to make moral judgments, have highlighted how technology will often lead to more adverse…

Artificial Intelligence · Computer Science 2022-03-22 Saif M. Mohammad