中文

GPIC:巨型宽松图像语料库用于视觉生成

计算机视觉与模式识别 2026-05-29 v1 人工智能

摘要

研究可扩展的视觉生成建模方法需要大型、可访问且稳定的数据集。我们引入 GPIC,这一约为 28 万亿像素的巨型宽松图像语料库。GPIC 包含来自互联网的多样化图像,由最先进的视觉语言模型标注,包括 100M 训练样本、20 万验证样本和 100 万测试样本。此外,所有 GPIC 图像均采用宽松许可证,可用于研究和商业用途。GPIC 经过安全筛选、去重,并在 Hugging Face 上集中托管。我们提供针对 GPIC 的生成式建模基准协议。最后,我们在 GPIC 上提供了像素空间流匹配的参考基线。我们的数据集、基准和模型均可在 https://huggingface.co/datasets/stanford-vision-lab/gpic 获取。评估工具包和代码可在 https://gpic.stanford.edu 获取。

关键词

引用

@article{arxiv.2605.30341,
  title  = {GPIC: A Giant Permissive Image Corpus for Visual Generation},
  author = {Keshigeyan Chandrasegaran and Kyle Sargent and Suchir Agarwal and Michael Jang and Michael Poli and Juan Carlos Niebles and Justin Johnson and Jiajun Wu and Li Fei-Fei},
  journal= {arXiv preprint arXiv:2605.30341},
  year   = {2026}
}

备注

25 pages; Dataset: https://huggingface.co/datasets/stanford-vision-lab/giant-permissive-image-corpus; Project website: https://gpic.stanford.edu