English
Related papers

Related papers: The State of Documentation Practices of Third-part…

200 papers

Background. The development of empirical studies in software engineering mainly relies on the data available on code hosting platforms, being GitHub the most representative. Nevertheless, in the last years, the emergence of Machine Learning…

Software Engineering · Computer Science 2023-07-28 Adem Ait , Javier Luis Cánovas Izquierdo , Jordi Cabot

Trained machine learning models are increasingly used to perform high-impact tasks in areas such as law enforcement, medicine, education, and employment. In order to clarify the intended use cases of machine learning models and minimize…

Deep learning models for natural language processing (NLP) are increasingly adopted and deployed by analysts without formal training in NLP or machine learning (ML). However, the documentation intended to convey the model's details and…

Human-Computer Interaction · Computer Science 2022-05-09 Anamaria Crisan , Margaret Drouhard , Jesse Vig , Nazneen Rajani

Developers are sharing pre-trained Machine Learning (ML) models through a variety of model sharing platforms, such as Hugging Face, in an effort to make ML development more collaborative. To share the models, they must first be serialized.…

Software Engineering · Computer Science 2025-01-07 Beatrice Casey , Kaia Damian , Andrew Cotaj , Joanna C. S. Santos

The development of machine learning (ML) techniques has led to ample opportunities for developers to develop and deploy their own models. Hugging Face serves as an open source platform where developers can share and download other models in…

Cryptography and Security · Computer Science 2024-10-08 Beatrice Casey , Joanna C. S. Santos , Mehdi Mirakhorli

The ubiquity of large-scale Pre-Trained Models (PTMs) is on the rise, sparking interest in model hubs, and dedicated platforms for hosting PTMs. Despite this trend, a comprehensive exploration of the challenges that users encounter and how…

Software Engineering · Computer Science 2024-01-25 Mina Taraghi , Gianolli Dorcelus , Armstrong Foundjem , Florian Tambon , Foutse Khomh

As machine learning (ML) becomes an integral part of high-autonomy systems, it is critical to ensure the trustworthiness of learning-enabled software systems (LESS). Yet, the nondeterministic and run-time-defined semantics of ML complicate…

Software Engineering · Computer Science 2025-12-10 Nan Jia , Anita Raja , Raffi Khatchadourian

Model cards are the primary documentation framework for developers of artificial intelligence (AI) models to communicate critical information to their users. Those users are often developers themselves looking for relevant documentation to…

Software Engineering · Computer Science 2025-11-20 Tim Puhlfürß , Julia Butzke , Walid Maalej

We analyzed nearly 460,000 AI model cards from Hugging Face to examine how developers report risks. From these, we extracted around 3,000 unique risk mentions and built the \emph{AI Model Risk Catalog}. We compared these with risks…

Computers and Society · Computer Science 2025-08-26 Pooja S. B. Rao , Sanja Šćepanović , Dinesh Babu Jayagopi , Mauro Cherubini , Daniele Quercia

Studies of dataset development in machine learning call for greater attention to the data practices that make model development possible and shape its outcomes. Many argue that the adoption of theory and practices from archives and data…

Computers and Society · Computer Science 2024-05-07 Eshta Bhardwaj , Harshit Gujral , Siyi Wu , Ciara Zogheib , Tegan Maharaj , Christoph Becker

Background: The development of AI-enabled software heavily depends on AI model documentation, such as model cards, due to different domain expertise between software engineers and model developers. From an ethical standpoint, AI model…

Software Engineering · Computer Science 2024-07-04 Haoyu Gao , Mansooreh Zahedi , Christoph Treude , Sarita Rosenstock , Marc Cheong

Data practices shape research and practice on fairness in machine learning (fair ML). Critical data studies offer important reflections and critiques for the responsible advancement of the field by highlighting shortcomings and proposing…

Machine Learning · Computer Science 2024-06-21 Jan Simson , Alessandro Fabris , Christoph Kern

Datasets are central to training machine learning (ML) models. The ML community has recently made significant improvements to data stewardship and documentation practices across the model development life cycle. However, the act of…

Computers and Society · Computer Science 2022-05-11 Alexandra Sasha Luccioni , Frances Corry , Hamsini Sridharan , Mike Ananny , Jason Schultz , Kate Crawford

Machine learning datasets have elicited concerns about privacy, bias, and unethical applications, leading to the retraction of prominent datasets such as DukeMTMC, MS-Celeb-1M, and Tiny Images. In response, the machine learning community…

Machine Learning · Computer Science 2021-11-23 Kenny Peng , Arunesh Mathur , Arvind Narayanan

Background: Open-Source Pre-Trained Models (PTMs) and datasets provide extensive resources for various Machine Learning (ML) tasks, yet these resources lack a classification tailored to Software Engineering (SE) needs. Aims: We apply an…

Software Engineering · Computer Science 2024-11-15 Alexandra González , Xavier Franch , David Lo , Silverio Martínez-Fernández

Recent advances in Artificial Intelligence, especially in Machine Learning (ML), have brought applications previously considered as science fiction (e.g., virtual personal assistants and autonomous cars) into the reach of millions of…

Software Engineering · Computer Science 2020-02-27 Minke Xiu , Zhen Ming , Jiang , Bram Adams

This article presents the current state of ML-security and of the documentation of ML-based systems, models and datasets in research and practice based on an extensive review of the existing literature. It shows a generally low awareness of…

Cryptography and Security · Computer Science 2025-07-17 Cara Ellen Appel

The availability of vast amounts of publicly accessible data of source code and the advances in modern language models, coupled with increasing computational resources, have led to a remarkable surge in the development of large language…

Software Engineering · Computer Science 2024-10-01 Zhou Yang , Jieke Shi , Premkumar Devanbu , David Lo

Empirically validating new 3D-printing related algorithms and implementations requires testing data representative of inputs encountered \emph{in the wild}. An ideal benchmarking dataset should not only draw from the same distribution of…

Graphics · Computer Science 2016-07-05 Qingnan Zhou , Alec Jacobson

Since late 2022, Large Language Models (LLMs) have become very prominent with LLMs like ChatGPT and Bard receiving millions of users. Hundreds of new LLMs are announced each week, many of which are deposited to Hugging Face, a repository of…

Digital Libraries · Computer Science 2023-07-20 Sarah Gao , Andrew Kean Gao