GPIC:巨型宽松图像语料库用于视觉生成
计算机视觉与模式识别
2026-05-29 v1 人工智能
摘要
研究可扩展的视觉生成建模方法需要大型、可访问且稳定的数据集。我们引入 GPIC,这一约为 28 万亿像素的巨型宽松图像语料库。GPIC 包含来自互联网的多样化图像,由最先进的视觉语言模型标注,包括 100M 训练样本、20 万验证样本和 100 万测试样本。此外,所有 GPIC 图像均采用宽松许可证,可用于研究和商业用途。GPIC 经过安全筛选、去重,并在 Hugging Face 上集中托管。我们提供针对 GPIC 的生成式建模基准协议。最后,我们在 GPIC 上提供了像素空间流匹配的参考基线。我们的数据集、基准和模型均可在 https://huggingface.co/datasets/stanford-vision-lab/gpic 获取。评估工具包和代码可在 https://gpic.stanford.edu 获取。
关键词
引用
@article{arxiv.2605.30341,
title = {GPIC: A Giant Permissive Image Corpus for Visual Generation},
author = {Keshigeyan Chandrasegaran and Kyle Sargent and Suchir Agarwal and Michael Jang and Michael Poli and Juan Carlos Niebles and Justin Johnson and Jiajun Wu and Li Fei-Fei},
journal= {arXiv preprint arXiv:2605.30341},
year = {2026}
}
备注
25 pages; Dataset: https://huggingface.co/datasets/stanford-vision-lab/giant-permissive-image-corpus; Project website: https://gpic.stanford.edu