中文

GPT-4o图像生成能力的实证研究

计算机视觉与模式识别 2025-04-14 v2

摘要

图像生成领域经历了快速演变,从早期的GAN方法到扩散模型,最新方法则致力于构建统一的生成架构,以跨越理解和生成任务。近期进展,尤其是GPT-4o,已展示了高保真度多模态生成的可行性,但其架构设计仍然神秘且未公开。这促使我们思考:文本和图像生成是否已经成功集成到统一框架中?本工作对GPT-4o的图像生成能力进行实证研究,将其与领先的开源和商业模型进行基准测试。我们的评估涵盖四个主要类别,包括文本到图像、图像到图像、图像到3D和图像到X生成,共计超过20项任务。我们的分析揭示了GPT-4o在不同设置下的优势与局限,并将其置于更广泛的生成模型演变脉络中。通过这一调查,我们识别了未来统一生成模型的有前景的方向,强调架构设计和数据规模的作用。有关高清PDF版本,请参阅GitHub上的链接:\href{https://github.com/Ephemeral182/Empirical-Study-of-GPT-4o-Image-Gen}{https://github.com/Ephemeral182/Empirical-Study-of-GPT-4o-Image-Gen}。

关键词

引用

@article{arxiv.2504.05979,
  title  = {An Empirical Study of GPT-4o Image Generation Capabilities},
  author = {Sixiang Chen and Jinbin Bai and Zhuoran Zhao and Tian Ye and Qingyu Shi and Donghao Zhou and Wenhao Chai and Xin Lin and Jianzong Wu and Chao Tang and Shilin Xu and Tao Zhang and Haobo Yuan and Yikang Zhou and Wei Chow and Linfeng Li and Xiangtai Li and Lei Zhu and Lu Qi},
  journal= {arXiv preprint arXiv:2504.05979},
  year   = {2025}
}