English

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs

Computer Vision and Pattern Recognition 2025-06-24 v1

Abstract

The integration of high-resolution image features in modern multimodal large language models has demonstrated significant improvements in fine-grained visual understanding tasks, achieving high performance across multiple benchmarks. Since these features are obtained from large image encoders like ViT, they come with a significant increase in computational costs due to multiple calls to these encoders. In this work, we first develop an intuition for feature upsampling as a natural extension of high-resolution feature generation. Through extensive experiments and ablations, we demonstrate how a shallow feature enricher can achieve competitive results with tremendous reductions in training and inference time as well as computational cost, with upto 1.5x saving in FLOPs.

Keywords

Cite

@article{arxiv.2506.17608,
  title  = {HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs},
  author = {Nikitha SR and Aradhya Neeraj Mathur and Tarun Ram Menta and Rishabh Jain and Mausoom Sarkar},
  journal= {arXiv preprint arXiv:2506.17608},
  year   = {2025}
}

Comments

Accepted in CVPR 2025 Workshop on What's Next in Multimodal Foundational Models