English

xGen-MM (BLIP-3): A Family of Open Large Multimodal Models

Computer Vision and Pattern Recognition 2025-09-18 v4 Artificial Intelligence Computation and Language

Abstract

This paper introduces BLIP-3, an open framework for developing Large Multimodal Models (LMMs). The framework comprises meticulously curated datasets, a training recipe, model architectures, and a resulting suite of LMMs. We release 4B and 14B models, including both the pre-trained base model and the instruction fine-tuned ones. Our models undergo rigorous evaluation across a range of tasks, including both single and multi-image benchmarks. Our models demonstrate competitive performance among open-source LMMs with similar model sizes. Our resulting LMMs demonstrate competitive performance among open-source LMMs with similar model sizes, with the ability to comprehend interleaved image-text inputs. Our training code, models, and all datasets used in this work, including the three largescale datasets we create and the preprocessed ones, will be open-sourced to better support the research community.

Keywords

Cite

@article{arxiv.2408.08872,
  title  = {xGen-MM (BLIP-3): A Family of Open Large Multimodal Models},
  author = {Le Xue and Manli Shu and Anas Awadalla and Jun Wang and An Yan and Senthil Purushwalkam and Honglu Zhou and Viraj Prabhu and Yutong Dai and Michael S Ryoo and Shrikant Kendre and Jieyu Zhang and Shaoyen Tseng and Gustavo A Lujan-Moreno and Matthew L Olson and Musashi Hinck and David Cobbley and Vasudev Lal and Can Qin and Shu Zhang and Chia-Chih Chen and Ning Yu and Juntao Tan and Tulika Manoj Awalgaonkar and Shelby Heinecke and Huan Wang and Yejin Choi and Ludwig Schmidt and Zeyuan Chen and Silvio Savarese and Juan Carlos Niebles and Caiming Xiong and Ran Xu},
  journal= {arXiv preprint arXiv:2408.08872},
  year   = {2025}
}