English
Related papers

Related papers: M4CXR: Exploring Multi-task Potentials of Multi-mo…

200 papers

Vision-language pretraining has advanced image-text alignment, yet progress in radiology remains constrained by the heterogeneity of clinical reports, including abbreviations, impression-only notes, and stylistic variability. Unlike…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Hanbin Ko , Gihun Cho , Inhyeok Baek , Donguk Kim , Joonbeom Koo , Changi Kim , Dongheon Lee , Chang Min Park

Multimodal Large Language Models (MLLMs) have shown success in various general image processing tasks, yet their application in medical imaging is nascent, lacking tailored models. This study investigates the potential of MLLMs in improving…

Computer Vision and Pattern Recognition · Computer Science 2024-05-13 Ling Yang , Zhanyu Wang , Zhenghao Chen , Xinyu Liang , Luping Zhou

Automated interpretation of chest X-rays (CXR) is a critical task with the potential to significantly improve clinical workflow and patient care. While recent advances in multimodal foundation models have shown promise, effectively…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Alexander Davis , Rafael Souza , Jia-Hao Lim

The rapid advancements in large language models (LLMs) have unlocked their potential for multimodal tasks, where text and visual data are processed jointly. However, applying LLMs to medical imaging, particularly for chest X-rays (CXR),…

Image and Video Processing · Electrical Eng. & Systems 2025-02-11 Nicholas Evans , Stephen Baker , Miles Reed

Following the impressive development of LLMs, vision-language alignment in LLMs is actively being researched to enable multimodal reasoning and visual IO. This direction of research is particularly relevant to medical imaging because…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Suhyeon Lee , Won Jun Kim , Jinho Chang , Jong Chul Ye

Longitudinal chest X-ray (CXR) interpretation requires reasoning over disease evolution across multiple patient visits, yet most existing medical VQA benchmarks focus on single images or short-horizon image pairs. We introduce MI-CXR, a…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Sunghwan Steve Cho , Yunseok Han , Jaeyoung Do

Artificial intelligence (AI)-based chest X-ray (CXR) interpretation assistants have demonstrated significant progress and are increasingly being applied in clinical settings. However, contemporary medical AI models often adhere to a…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Jinquan Guan , Qi Chen , Lizhou Liang , Yuhang Liu , Vu Minh Hieu Phan , Minh-Son To , Jian Chen , Yutong Xie

The widespread use of chest X-rays (CXRs), coupled with a shortage of radiologists, has driven growing interest in automated CXR analysis and AI-assisted reporting. While existing vision-language models (VLMs) show promise in specific tasks…

Chest X-ray (CXR) imaging is one of the most widely used diagnostic modalities in clinical practice, encompassing a broad spectrum of diagnostic tasks. Recent advancements have seen the extensive application of reasoning-based multimodal…

Chest X-rays (CXRs) are the most frequently performed imaging examinations in clinical settings. Recent advancements in Large Multimodal Models (LMMs) have enabled automated CXR interpretation, enhancing diagnostic accuracy and efficiency.…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Qingqiu Li , Zihang Cui , Seongsu Bae , Jilan Xu , Runtian Yuan , Yuejie Zhang , Rui Feng , Quanli Shen , Xiaobo Zhang , Junjun He , Shujun Wang

Purpose: This study aimed to develop an open-source multimodal large language model (CXR-LLAVA) for interpreting chest X-ray images (CXRs), leveraging recent advances in large language models (LLMs) to potentially replicate the image…

Computation and Language · Computer Science 2024-01-17 Seowoo Lee , Jiwon Youn , Hyungjin Kim , Mansu Kim , Soon Ho Yoon

Significant methodological strides have been made toward Chest X-ray (CXR) understanding via modern vision-language models (VLMs), demonstrating impressive Visual Question Answering (VQA) and CXR report generation abilities. However,…

Artificial Intelligence · Computer Science 2024-04-01 Seil Kang , Donghyun Kim , Junhyeok Kim , Hyo Kyung Lee , Seong Jae Hwang

The latest breakthroughs in large vision-language models, such as Bard and GPT-4, have showcased extraordinary abilities in performing a wide range of tasks. Such models are trained on massive datasets comprising billions of public…

Computer Vision and Pattern Recognition · Computer Science 2025-05-08 Omkar Thawakar , Abdelrahman Shaker , Sahal Shaji Mullappilly , Hisham Cholakkal , Rao Muhammad Anwer , Salman Khan , Jorma Laaksonen , Fahad Shahbaz Khan

Recent advances in reasoning-enhanced large language models (LLMs) and multimodal LLMs (MLLMs) have significantly improved performance in complex tasks, yet medical AI models often overlook the structured reasoning processes inherent in…

Artificial Intelligence · Computer Science 2025-05-22 Ziqing Fan , Cheng Liang , Chaoyi Wu , Ya Zhang , Yanfeng Wang , Weidi Xie

Chest X-ray (CXR) plays a pivotal role in clinical diagnosis, and a variety of task-specific and foundation models have been developed for automatic CXR interpretation. However, these models often struggle to adapt to new diagnostic tasks…

Artificial Intelligence · Computer Science 2025-10-27 Jinhui Lou , Yan Yang , Zhou Yu , Zhenqi Fu , Weidong Han , Qingming Huang , Jun Yu

We introduce a radiology-focused visual language model designed to generate radiology reports from chest X-rays. Building on previous findings that large language models (LLMs) can acquire multimodal capabilities when aligned with…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Xi Zhang , Zaiqiao Meng , Jake Lever , Edmond S. L. Ho

3D medical image analysis is of great importance in disease diagnosis and treatment. Recently, multimodal large language models (MLLMs) have exhibited robust perceptual capacity, strong cross-modal alignment, and promising generalizability.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Yang Yu , Dunyuan Xu , Yaoqian Li , Xiaomeng Li , Jinpeng Li , Pheng-Ann Heng

Recently, Multimodal Large Language Model (MLLM) represented by GPT-4V has been a new rising research hotspot, which uses powerful Large Language Models (LLMs) as a brain to perform multimodal tasks. The surprising emergent capabilities of…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Shukang Yin , Chaoyou Fu , Sirui Zhao , Ke Li , Xing Sun , Tong Xu , Enhong Chen

The rapid advancement of Large Language Models (LLMs) has opened new possibilities in Multi-Robot Systems (MRS), enabling enhanced communication, task allocation and planning, and human-robot interaction. Unlike traditional single-robot and…

Robotics · Computer Science 2026-05-05 Peihan Li , Zijian An , Shams Abrar , Lifeng Zhou

Extremely large-scale massive multiple-input multiple-output (XL-MIMO) is a key enabler for sixth-generation (6G) networks, offering massive spatial degrees of freedom. Despite these advantages, the coexistence of near-field and far-field…

Machine Learning · Computer Science 2025-12-11 Renbin Li , Shuangshuang Li , Peihao Dong
‹ Prev 1 2 3 10 Next ›