中文

OmniCaptioner:一条 Caption 征服万物

计算机视觉与模式识别 2025-06-03 v3 计算与语言

摘要

我们提出 OmniCaptioner,一个用于生成细粒度文本描述的通用视觉描述框架,覆盖广泛的视觉领域。不同于仅限于特定图像类型(如自然图像或几何图形)的先前方法,我们的框架为描述自然图像、视觉文本(如海报、界面、教科书)以及结构化视觉内容(如文档、表格、图表)提供统一解决方案。通过将低级像素信息转换为语义丰富的文本表示,我们的框架搭建了视觉与文本模态之间的鸿沟。我们的结果突显了三个关键优势:(i)增强视觉推理,借助视觉模态的长上下文描述,特别是 DeepSeek-R1 系列模型,在多模态场景中能够有效推理;(ii)改进图像生成,详细描述改善了文本到图像生成和图像转换等任务;(iii)高效监督微调(SFT),实现更少数据即可实现更快收敛。我们认为 OminiCaptioner 的通用性和适应性为搭建语言与视觉模态之间的新视角提供了可能。

关键词

引用

@article{arxiv.2504.07089,
  title  = {OmniCaptioner: One Captioner to Rule Them All},
  author = {Yiting Lu and Jiakang Yuan and Zhen Li and Shitian Zhao and Qi Qin and Xinyue Li and Le Zhuo and Licheng Wen and Dongyang Liu and Yuewen Cao and Xiangchao Yan and Xin Li and Tianshuo Peng and Shufei Zhang and Botian Shi and Tao Chen and Zhibo Chen and Lei Bai and Peng Gao and Bo Zhang},
  journal= {arXiv preprint arXiv:2504.07089},
  year   = {2025}
}

备注

More visualizations on Homepage: https://alpha-innovator.github.io/OmniCaptioner-project-page and Official code: https://github.com/Alpha-Innovator/OmniCaptioner