中文

BLIP3-KALE:知识增强的大规模稠密描述

计算机视觉与模式识别 2024-11-13 v1 人工智能

摘要

我们推出 BLIP3-KALE,一个包含 2.18 亿图文对的数据集,填补了描述性合成标题与事实性网络规模替代文本之间的空白。KALE 通过网络规模的替代文本增强合成稠密图像标题,生成有事实依据的图像标题。我们的两阶段方法利用大型视觉-语言模型和语言模型创建知识增强标题,进而用于训练专用 VLM 以扩展数据集规模。我们在 KALE 上训练视觉-语言模型,并在视觉-语言任务上展示了性能提升。实验表明 KALE 对训练更强大、更知识化的多模态模型具有实用价值。我们在 https://huggingface.co/datasets/Salesforce/blip3-kale 发布 KALE 数据集。

关键词

引用

@article{arxiv.2411.07461,
  title  = {BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions},
  author = {Anas Awadalla and Le Xue and Manli Shu and An Yan and Jun Wang and Senthil Purushwalkam and Sheng Shen and Hannah Lee and Oscar Lo and Jae Sung Park and Etash Guha and Silvio Savarese and Ludwig Schmidt and Yejin Choi and Caiming Xiong and Ran Xu},
  journal= {arXiv preprint arXiv:2411.07461},
  year   = {2024}
}