第三方机器学习模型与数据集文档实践的现状
软件工程
2024-06-19 v1 机器学习
摘要
模型商店提供第三方机器学习模型和数据集,便于项目集成,减少编码工作量。人们可能希望在文档中找到这些模型和数据集的详细规范,并利用模型卡和数据集卡等文档标准。在本研究中,我们使用统计分析和混合卡片分类法,评估了当前最大的模型商店之一——Hugging Face(HF)中模型卡和数据集卡文档的实践现状。我们的发现表明,只有21,902个模型(39.62%)和1,925个数据集(28.48%)具有文档。此外,我们观察到机器学习模型和数据集的伦理与透明度相关文档存在不一致性。
引用
@article{arxiv.2312.15058,
title = {The State of Documentation Practices of Third-party Machine Learning Models and Datasets},
author = {Ernesto Lang Oreamuno and Rohan Faiyaz Khan and Abdul Ali Bangash and Catherine Stinson and Bram Adams},
journal= {arXiv preprint arXiv:2312.15058},
year = {2024}
}
备注
7 pages, 4 figures, IEEESoftware format