English
Related papers

Related papers: The Data Provenance Initiative: A Large Scale Audi…

200 papers

The advent of Generative Artificial Intelligence (GenAI) models, including GitHub Copilot, OpenAI GPT, and Stable Diffusion, has revolutionized content creation, enabling non-professionals to produce high-quality content across various…

Computer Vision and Pattern Recognition · Computer Science 2024-05-08 Uri Hacohen , Adi Haviv , Shahar Sarfaty , Bruria Friedman , Niva Elkin-Koren , Roi Livni , Amit H Bermano

The integration of generative artificial intelligence (GenAI) and large language models (LLMs) into scientific research and higher education presents a paradigm shift, offering revolutionizing opportunities while simultaneously raising…

Digital Libraries · Computer Science 2026-01-21 Dmitry Kochetkov

The utilization of artificial intelligence (AI) applications has experienced tremendous growth in recent years, bringing forth numerous benefits and conveniences. However, this expansion has also provoked ethical concerns, such as privacy…

Advances in AI-generated content have led to wide adoption of large language models, diffusion-based visual generators, and synthetic audio tools. However, these developments raise critical concerns about misinformation, copyright…

Computation and Language · Computer Science 2025-09-30 Lele Cao

As data-driven systems are increasingly deployed at scale, ethical concerns have arisen around unfair and discriminatory outcomes for historically marginalized groups that are underrepresented in training data. In response, work around AI…

Human-Computer Interaction · Computer Science 2022-09-21 Rie Kamikubo , Lining Wang , Crystal Marte , Amnah Mahmood , Hernisa Kacorri

High availability of data is responsible for the current trends in Artificial Intelligence (AI) and Machine Learning (ML). However, high-grade datasets are reluctantly shared between actors because of lacking trust and fear of losing…

Cryptography and Security · Computer Science 2020-02-26 Philipp Lüthi , Thibault Gagnaux , Marcel Gygli

Open datasets play a crucial role in three research domains that intersect data science and education: learning analytics, educational data mining, and artificial intelligence in education. Researchers in these domains apply computational…

Computers and Society · Computer Science 2026-04-14 Valdemar Švábenský , Brendan Flanagan , Erwin Daniel López Zapata , Atsushi Shimada

Generative Artificial Intelligence (GAI) has experienced exponential growth in recent years, partly facilitated by the abundance of large-scale open-source datasets. These datasets are often built using unrestricted and opaque data…

Artificial Intelligence · Computer Science 2025-12-29 Matyas Bohacek , Ignacio Vilanova Echavarri

Despite the utility that Generative AI (GenAI) tools provide for tasks such as writing code, the use of these tools raises important legal questions and potential risks, particularly those associated with copyright law. As lawmakers and…

Datasets sourced from people with disabilities and older adults play an important role in innovation, benchmarking, and mitigating bias for both assistive and inclusive AI-infused applications. However, they are scarce. We conduct a…

Human-Computer Interaction · Computer Science 2021-08-25 Rie Kamikubo , Utkarsh Dwivedi , Hernisa Kacorri

Diffusion Models (DMs) benefit from large and diverse datasets for their training. Since this data is often scraped from the Internet without permission from the data owners, this raises concerns about copyright and intellectual property…

Machine Learning · Computer Science 2025-06-24 Jan Dubiński , Antoni Kowalczuk , Franziska Boenisch , Adam Dziedzic

Generative audio models are rapidly advancing in both capabilities and public utilization -- several powerful generative audio models have readily available open weights, and some tech companies have released high quality generative audio…

Data-centric AI is at the center of a fundamental shift in software engineering where machine learning becomes the new software, powered by big data and computing infrastructure. Here software engineering needs to be re-thought where data…

Machine Learning · Computer Science 2022-12-27 Steven Euijong Whang , Yuji Roh , Hwanjun Song , Jae-Gil Lee

Dataset license compliance is a critical yet complex aspect of developing commercial AI products, particularly with the increasing use of publicly available datasets. Ambiguities in dataset licenses pose significant legal risks, making it…

Software Engineering · Computer Science 2026-05-08 Jingwen Tan , Gopi Krishnan Rajbahadur , Zi Li , Xiangfu Song , Jianshan Lin , Dan Li , Zibin Zheng , Ahmed E. Hassan

The proliferation of generative AI systems has created new challenges for the Free and Open Source Software (FOSS) community, particularly regarding how traditional copyleft principles should apply when open source code is used to train AI…

Computers and Society · Computer Science 2026-02-09 Grant Shanklin , Emmie Hine , Claudio Novelli , Tyler Schroder , Luciano Floridi

Context: When software is released publicly, it is common to include with it either the full text of the license or licenses under which it is published, or a detailed reference to them. Therefore public licenses, including FOSS (free, open…

Software Engineering · Computer Science 2023-08-23 Jesús M. González-Barahona , Sergio Montes-Leon , Gregorio Robles , Stefano Zacchiroli

Pre-training, which utilizes extensive and varied datasets, is a critical factor in the success of Large Language Models (LLMs) across numerous applications. However, the detailed makeup of these datasets is often not disclosed, leading to…

Cryptography and Security · Computer Science 2024-01-02 Haodong Li , Gelei Deng , Yi Liu , Kailong Wang , Yuekang Li , Tianwei Zhang , Yang Liu , Guoai Xu , Guosheng Xu , Haoyu Wang

Our analysis of recent AI4H publications reveals that, despite a trend toward utilizing open datasets and sharing modeling code, 74% of AI4H papers still rely on private datasets or do not share their code. This is especially concerning in…

Computers and Society · Computer Science 2026-03-05 John Wu , Zhenbang Wu , Jimeng Sun

Successful data-driven science requires complex data engineering pipelines to clean, transform, and alter data in preparation for machine learning, and robust results can only be achieved when each step in the pipeline can be justified, and…

Databases · Computer Science 2024-04-08 Adriane Chapman , Luca Lauro , Paolo Missier , Riccardo Torlone

The growing adoption of artificial intelligence (AI) has amplified concerns about trustworthiness, including integrity, privacy, robustness, and bias. To assess and attribute these threats, we propose ConceptLens, a generic framework that…

Cryptography and Security · Computer Science 2025-07-17 Jiamin Chang , Haoyang Li , Hammond Pearce , Ruoxi Sun , Bo Li , Minhui Xue