English

Can Large Language Models Challenge CNNs in Medical Image Analysis?

Image and Video Processing 2025-06-04 v2 Artificial Intelligence Computer Vision and Pattern Recognition

Abstract

This study presents a multimodal AI framework designed for precisely classifying medical diagnostic images. Utilizing publicly available datasets, the proposed system compares the strengths of convolutional neural networks (CNNs) and different large language models (LLMs). This in-depth comparative analysis highlights key differences in diagnostic performance, execution efficiency, and environmental impacts. Model evaluation was based on accuracy, F1-score, average execution time, average energy consumption, and estimated CO2CO_2 emission. The findings indicate that although CNN-based models can outperform various multimodal techniques that incorporate both images and contextual information, applying additional filtering on top of LLMs can lead to substantial performance gains. These findings highlight the transformative potential of multimodal AI systems to enhance the reliability, efficiency, and scalability of medical diagnostics in clinical settings.

Keywords

Cite

@article{arxiv.2505.23503,
  title  = {Can Large Language Models Challenge CNNs in Medical Image Analysis?},
  author = {Shibbir Ahmed and Shahnewaz Karim Sakib and Anindya Bijoy Das},
  journal= {arXiv preprint arXiv:2505.23503},
  year   = {2025}
}
R2 v1 2026-07-01T02:48:32.061Z