Public Domain 12M:具有新颖治理机制的高美学图像-文本数据集
人工智能
2024-10-31 v1
摘要
我们展示了Public Domain 12M(PD12M),一个包含1240万张高质量公有领域和CC0授权图像及合成标注的数据集,旨在用于训练文本到图像模型。PD12M是迄今为止最大的公有领域图像-文本数据集,规模足以训练基础模型,同时最大限度地减少版权问题。通过Source.Plus平台,我们还引入了新颖的、社区驱动的数据集治理机制,以减少危害并支持长期的可复现性。
引用
@article{arxiv.2410.23144,
title = {Public Domain 12M: A Highly Aesthetic Image-Text Dataset with Novel Governance Mechanisms},
author = {Jordan Meyer and Nick Padgett and Cullen Miller and Laura Exline},
journal= {arXiv preprint arXiv:2410.23144},
year = {2024}
}
备注
Project Page: https://source.plus/pd12m