English
Related papers

Related papers: Argus: Benchmarking and Enhancing Vision-Language …

200 papers

Multimodal Large Language Models (MLLMs) have shown strong potential for radiology report generation, yet their clinical translation is hindered by architectural heterogeneity and the prevalence of factual hallucinations. Standard…

Machine Learning · Computer Science 2026-01-13 Kun Zhao , Siyuan Dai , Pan Wang , Jifeng Song , Hui Ji , Chenghua Lin , Liang Zhan , Haoteng Tang

This review provides a systematic analysis of comprehensive survey of 3D object detection with vision-language models(VLMs) , a rapidly advancing area at the intersection of 3D vision and multimodal AI. By examining over 100 research…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Ranjan Sapkota , Konstantinos I Roumeliotis , Rahul Harsha Cheppally , Marco Flores Calero , Manoj Karkee

Medical image analysis is essential to clinical diagnosis and treatment, which is increasingly supported by multi-modal large language models (MLLMs). However, previous research has primarily focused on 2D medical images, leaving 3D images…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Fan Bai , Yuxin Du , Tiejun Huang , Max Q. -H. Meng , Bo Zhao

Automated radiology report generation from 3D CT volumes often suffers from incomplete pathology coverage. We provide empirical evidence that this limitation stems from a representational bottleneck: contrastive 3D CT embeddings encode…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Renjie Liang , Yiling Ma , Yang Xing , Zhengkang Fan , Jinqian Pan , Chengkun Sun , Li Li , Kuang Gong , Jie Xu

Open-vocabulary 3D Scene Graph (3DSG) can enhance various downstream tasks in robotics by leveraging structured semantic representations, yet current 3DSG construction methods suffer from semantic inconsistencies caused by noisy cross-image…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Yue Chang , Rufeng Chen , Zhaofan Zhang , Yi Chen , Yifan Tian , Sihong Xie

Advancements in generative Artificial Intelligence (AI) hold great promise for automating radiology workflows, yet challenges in interpretability and reliability hinder clinical adoption. This paper presents an automated radiology report…

Positron emission tomography (PET) is a cornerstone of modern oncologic and neurologic imaging, distinguished by its unique ability to illuminate dynamic metabolic processes that transcend the anatomical focus of traditional imaging…

Image and Video Processing · Electrical Eng. & Systems 2025-11-13 Yichi Zhang , Wenbo Zhang , Zehui Ling , Gang Feng , Sisi Peng , Deshu Chen , Yuchen Liu , Hongwei Zhang , Shuqi Wang , Lanlan Li , Limei Han , Yuan Cheng , Zixin Hu , Yuan Qi , Le Xue

Developing 3D vision-language models with robust clinical reasoning remains a challenge due to the inherent complexity of volumetric medical imaging, the tendency of models to overfit superficial report patterns, and the lack of…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Haoran Lai , Zihang Jiang , Kun Zhang , Qingsong Yao , Rongsheng Wang , Zhiyang He , Xiaodong Tao , Wei Wei , Shaohua Kevin Zhou

Vision Large Language Models (VLMs) combine visual understanding with natural language processing, enabling tasks like image captioning, visual question answering, and video analysis. While VLMs show impressive capabilities across domains…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Ahmed Sharshar , Latif U. Khan , Waseem Ullah , Mohsen Guizani

Automated radiology report generation is essential in clinical practice. However, diagnosing radiological images typically requires physicians 5-10 minutes, resulting in a waste of valuable healthcare resources. Existing studies have not…

Multimedia · Computer Science 2025-09-16 Jing Xiao , Hongfei Liu , Ruiqi Dong , Jimin Liu , Haoyong Yu

The world faces a shortage of radiologists, leading to longer treatment times and increased stress, negatively impacting patient safety and workforce morale. Integrating artificial intelligence to interpret radiographic images and generate…

Image and Video Processing · Electrical Eng. & Systems 2024-06-19 Marijn Borghouts

Medical Visual Question Answering (Med-VQA) holds significant potential for clinical decision support, yet existing efforts primarily focus on 2D imaging with limited task diversity. This paper presents 3D-RAD, a large-scale dataset…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Xiaotang Gai , Jiaxiang Liu , Yichen Li , Zijie Meng , Jian Wu , Zuozhu Liu

Medical report generation from imaging data remains a challenging task in clinical practice. While large language models (LLMs) show great promise in addressing this challenge, their effective integration with medical imaging data still…

Computer Vision and Pattern Recognition · Computer Science 2025-06-19 Chunlei Li , Jingyang Hou , Yilei Shi , Jingliang Hu , Xiao Xiang Zhu , Lichao Mou

Writing radiology reports from medical images requires a high level of domain expertise. It is time-consuming even for trained radiologists and can be error-prone for inexperienced radiologists. It would be appealing to automate this task…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Yuzhe Lu , Sungmin Hong , Yash Shah , Panpan Xu

Accurate radio-frequency (RF) material parameters are essential for electromagnetic digital twins in 6G systems, yet gradient-based inverse ray tracing (RT) remains sensitive to initialization and costly under limited measurements. This…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Zerui Kang , Yishen Lim , Zhouyou Gu , Seung-Woo Ko , Tony Q. S. Quek , Jihong Park

Radiology report generation (RRG) for diagnostic images, such as chest X-rays, plays a pivotal role in both clinical practice and AI. Traditional free-text reports suffer from redundancy and inconsistent language, complicating the…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Yingshu Li , Yunyi Liu , Zhanyu Wang , Xinyu Liang , Lingqiao Liu , Lei Wang , Luping Zhou

Vision-language models (VLMs) have shown remarkable progress in offline tasks such as image captioning and video question answering. However, real-time interactive environments impose new demands on VLMs, requiring them to generate…

Computer Vision and Pattern Recognition · Computer Science 2025-05-19 Keunwoo Peter Yu , Joyce Chai

Vision-Language Models (VLMs) trained on web-scale corpora excel at natural image tasks and are increasingly repurposed for healthcare; however, their competence in medical tasks remains underexplored. We present a comprehensive evaluation…

Computer Vision and Pattern Recognition · Computer Science 2025-07-21 Che Liu , Jiazhen Pan , Weixiang Shen , Wenjia Bai , Daniel Rueckert , Rossella Arcucci

Conversational AI tools that can generate and discuss clinically correct radiology reports for a given medical image have the potential to transform radiology. Such a human-in-the-loop radiology assistant could facilitate a collaborative…

Computer Vision and Pattern Recognition · Computer Science 2025-05-08 Chantal Pellegrini , Ege Özsoy , Benjamin Busam , Nassir Navab , Matthias Keicher

Current vision-language models (VLMs) in medicine are primarily designed for categorical question answering (e.g., "Is this normal or abnormal?") or qualitative descriptive tasks. However, clinical decision-making often relies on…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Yongcheng Yao , Yongshuo Zong , Raman Dutt , Yongxin Yang , Sotirios A Tsaftaris , Timothy Hospedales
‹ Prev 1 4 5 6 7 8 10 Next ›