English

Focus on What Matters: Enhancing Medical Vision-Language Models with Automatic Attention Alignment Tuning

Computer Vision and Pattern Recognition 2025-05-27 v1

Abstract

Medical Large Vision-Language Models (Med-LVLMs) often exhibit suboptimal attention distribution on visual inputs, leading to hallucinated or inaccurate outputs. Existing mitigation methods primarily rely on inference-time interventions, which are limited in attention adaptation or require additional supervision. To address this, we propose A3^3Tune, a novel fine-tuning framework for Automatic Attention Alignment Tuning. A3^3Tune leverages zero-shot weak labels from SAM, refines them into prompt-aware labels using BioMedCLIP, and then selectively modifies visually-critical attention heads to improve alignment while minimizing interference. Additionally, we introduce a A3^3MoE module, enabling adaptive parameter selection for attention tuning across diverse prompts and images. Extensive experiments on medical VQA and report generation benchmarks show that A3^3Tune outperforms state-of-the-art baselines, achieving enhanced attention distributions and performance in Med-LVLMs.

Keywords

Cite

@article{arxiv.2505.18503,
  title  = {Focus on What Matters: Enhancing Medical Vision-Language Models with Automatic Attention Alignment Tuning},
  author = {Aofei Chang and Le Huang and Alex James Boyd and Parminder Bhatia and Taha Kass-Hout and Cao Xiao and Fenglong Ma},
  journal= {arXiv preprint arXiv:2505.18503},
  year   = {2025}
}

Comments

Accepted to ACL2025 (main)