English

Benchmarking Large Multimodal Models against Common Corruptions

Machine Learning 2024-01-23 v1 Computation and Language Cryptography and Security Computer Vision and Pattern Recognition Multimedia

Abstract

This technical report aims to fill a deficiency in the assessment of large multimodal models (LMMs) by specifically examining the self-consistency of their outputs when subjected to common corruptions. We investigate the cross-modal interactions between text, image, and speech, encompassing four essential generation tasks: text-to-image, image-to-text, text-to-speech, and speech-to-text. We create a comprehensive benchmark, named MMCBench, that covers more than 100 popular LMMs (totally over 150 model checkpoints). A thorough evaluation under common corruptions is critical for practical deployment and facilitates a better understanding of the reliability of cutting-edge LMMs. The benchmarking code is available at https://github.com/sail-sg/MMCBench

Keywords

Cite

@article{arxiv.2401.11943,
  title  = {Benchmarking Large Multimodal Models against Common Corruptions},
  author = {Jiawei Zhang and Tianyu Pang and Chao Du and Yi Ren and Bo Li and Min Lin},
  journal= {arXiv preprint arXiv:2401.11943},
  year   = {2024}
}

Comments

Technical report

R2 v1 2026-06-28T14:23:30.565Z