English

RoiMAM: Region-of-Interest Medical Attention Model for Efficient Vision-Language Understanding

Computer Vision and Pattern Recognition 2026-05-18 v1

Abstract

Vision-Language Models (VLMs) facilitate medical visual question answering (MedVQA) by jointly interpreting images and text. However, existing models typically depend on large architectures and closed-set answers, which limits their efficiency and potential clinical applicability. To overcome these shortcomings, we introduce RoiMAM, an efficient VLM. It integrates a training-free ROI Generation Module with Semantic Selective Suppression to focus on lesion-relevant regions, alongside a Text Prompt Enhancer module that provides modality-specific context without introducing training parameters. Compared to the widely used MedVInT-TD model, our design achieves efficient and accurate diagnosis at less than 20\% of the model size, while improving accuracy by approximately 2% on SLAKE and 4.6% on PMC-VQA.

Keywords

Cite

@article{arxiv.2605.15561,
  title  = {RoiMAM: Region-of-Interest Medical Attention Model for Efficient Vision-Language Understanding},
  author = {Jiayan Yang and Zhuoyu Wu and Wenqi Fang},
  journal= {arXiv preprint arXiv:2605.15561},
  year   = {2026}
}

Comments

under revision