English
Related papers

Related papers: Knowledge-enhanced Visual-Language Pre-training on…

200 papers

In emergency departments, rural hospitals, or clinics in less developed regions, clinicians often lack fast image analysis by trained radiologists, which can have a detrimental effect on patients' healthcare. Large Language Models (LLMs)…

Artificial Intelligence · Computer Science 2024-09-11 David Bani-Harouni , Nassir Navab , Matthias Keicher

Large-scale biomedical vision-language models (VLMs) adapted on high-end imaging (e.g., CT) often fail to transfer to frontline low-end modalities (e.g., radiography), collapsing into modality-specific shortcuts. We propose K-MaT…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Jiajun Zeng , Shadi Albarqouni

Medical Visual Question Answering (Med-VQA) holds significant potential for clinical decision support, yet existing efforts primarily focus on 2D imaging with limited task diversity. This paper presents 3D-RAD, a large-scale dataset…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Xiaotang Gai , Jiaxiang Liu , Yichen Li , Zijie Meng , Jian Wu , Zuozhu Liu

Disease classification relying solely on imaging data attracts great interest in medical image analysis. Current models could be further improved, however, by also employing Electronic Health Records (EHRs), which contain rich information…

Image and Video Processing · Electrical Eng. & Systems 2021-03-22 Tom van Sonsbeek , Xiantong Zhen , Marcel Worring , Ling Shao

A large-scale image-text pair dataset has greatly contributed to the development of vision-language pre-training (VLP) models, which enable zero-shot or few-shot classification without costly annotation. However, in the medical domain, the…

Computer Vision and Pattern Recognition · Computer Science 2023-10-23 Kihyun You , Jawook Gu , Jiyeon Ham , Beomhee Park , Jiho Kim , Eun Kyoung Hong , Woonhyunk Baek , Byungseok Roh

Pre-trained vision-language models, e.g., CLIP, working with manually designed prompts have demonstrated great capacity of transfer learning. Recently, learnable prompts achieve state-of-the-art performance, which however are prone to…

Computer Vision and Pattern Recognition · Computer Science 2023-08-23 Baoshuo Kan , Teng Wang , Wenpeng Lu , Xiantong Zhen , Weili Guan , Feng Zheng

Large annotated datasets are essential for training robust Computer-Aided Diagnosis (CAD) models for breast cancer detection or risk prediction. However, acquiring such datasets with fine-detailed annotation is both costly and…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Shunjie-Fabian Zheng , Hyeonjun Lee , Thijs Kooi , Ali Diba

Medical visual question answering (VQA) is a challenging task that requires answering clinical questions of a given medical image, by taking consider of both visual and language information. However, due to the small scale of training data…

Computer Vision and Pattern Recognition · Computer Science 2023-07-12 Pengfei Li , Gang Liu , Jinlong He , Zixu Zhao , Shenjun Zhong

Pathology detection and delineation enables the automatic interpretation of medical scans such as chest X-rays while providing a high level of explainability to support radiologists in making informed decisions. However, annotating…

Computer Vision and Pattern Recognition · Computer Science 2023-09-07 Philip Müller , Felix Meissen , Johannes Brandt , Georgios Kaissis , Daniel Rueckert

Purpose: As visual inspection is an inherent process during radiological screening, the associated eye gaze data can provide valuable insights into relevant clinical decisions. As deep learning has become the state-of-the-art for…

Image and Video Processing · Electrical Eng. & Systems 2025-02-18 Zirui Qiu , Hassan Rivaz , Yiming Xiao

This study presents a computer-aided diagnosis (CAD) system to assist early detection of lung metastases during endobronchial ultrasound (EBUS) procedures, significantly reducing follow-up time and enabling timely treatment. Due to limited…

Image and Video Processing · Electrical Eng. & Systems 2025-05-15 Ching-Kai Lin , Di-Chun Wei , Yun-Chien Cheng

Recent advancements in multimodal models have significantly improved vision-language (VL) alignment in radiology. However, existing approaches struggle to effectively utilize complex radiology reports for learning and offer limited…

Computer Vision and Pattern Recognition · Computer Science 2025-11-07 Jonggwon Park , Byungmu Yoon , Soobum Kim , Kyoyun Choi

Breast cancer remains the most commonly diagnosed malignancy among women in the developed world. Early detection through mammography screening plays a pivotal role in reducing mortality rates. While computer-aided diagnosis (CAD) systems…

Computer Vision and Pattern Recognition · Computer Science 2025-10-30 Shunjie-Fabian Zheng , Hyeonjun Lee , Thijs Kooi , Ali Diba

Lesion detection, symptom tracking, and visual explainability are central to real-world medical image analysis, yet current medical Vision-Language Models (VLMs) still lack mechanisms that translate their broad knowledge into clinically…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Woohyeon Park , Jaeik Kim , Sunghwan Steve Cho , Pa Hong , Wookyoung Jeong , Yoojin Nam , Namjoon Kim , Ginny Y. Wong , Ka Chun Cheung , Jaeyoung Do

To facilitate zero-shot generalization in taskoriented dialog, this paper proposes Language Models as Data (LAD). LAD is a paradigm for creating diverse and accurate synthetic data which conveys the necessary structural constraints and can…

Computation and Language · Computer Science 2022-08-01 Shikib Mehri , Yasemin Altun , Maxine Eskenazi

Automatic diagnosis (AD), a critical application of AI in healthcare, employs machine learning techniques to assist doctors in gathering patient symptom information for precise disease diagnosis. The Transformer-based method utilizes an…

Computation and Language · Computer Science 2023-07-18 Huimin Wang , Wai-Chung Kwan , Kam-Fai Wong , Yefeng Zheng

To contribute to automating the medical vision-language model, we propose a novel Chest-Xray Difference Visual Question Answering (VQA) task. Given a pair of main and reference images, this task attempts to answer several questions on both…

Computer Vision and Pattern Recognition · Computer Science 2024-08-29 Xinyue Hu , Lin Gu , Qiyuan An , Mengliang Zhang , Liangchen Liu , Kazuma Kobayashi , Tatsuya Harada , Ronald M. Summers , Yingying Zhu

The pandemic resulted in vast repositories of unstructured data, including radiology reports, due to increased medical examinations. Previous research on automated diagnosis of COVID-19 primarily focuses on X-ray images, despite their lower…

Vision-language foundation models have shown great promise in computational pathology but remain primarily data-driven, lacking explicit integration of medical knowledge. We introduce KEEP (KnowledgE-Enhanced Pathology), a foundation model…

Image and Video Processing · Electrical Eng. & Systems 2026-01-28 Xiao Zhou , Luoyi Sun , Dexuan He , Wenbin Guan , Ge Wang , Ruifen Wang , Lifeng Wang , Xiaojun Yuan , Xin Sun , Ya Zhang , Kun Sun , Yanfeng Wang , Weidi Xie

Lung cancer has been one of the major threats to human life for decades. Computer-aided diagnosis can help with early lung nodul detection and facilitate subsequent nodule characterization. Large Visual Language models (VLMs) have been…

Image and Video Processing · Electrical Eng. & Systems 2024-07-04 Furqan Shaukat , Syed Muhammad Anwar , Abhijeet Parida , Van Khanh Lam , Marius George Linguraru , Mubarak Shah