English
Related papers

Related papers: A multi-modal vision-language model for generaliza…

200 papers

Object detection is of paramount importance in biomedical image analysis, particularly for lesion identification. While current methodologies are proficient in identifying and pinpointing lesions, they often lack the precision needed to…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Zilin Chen , Shengnan Lu

While high-resolution pathology images lend themselves well to `data hungry' deep learning algorithms, obtaining exhaustive annotations on these images is a major challenge. In this paper, we propose a self-supervised CNN approach to…

Computer Vision and Pattern Recognition · Computer Science 2020-08-14 Navid Alemi Koohbanani , Balagopal Unnikrishnan , Syed Ali Khurram , Pavitra Krishnaswamy , Nasir Rajpoot

Noninvasive optical imaging modalities can probe patient's tissue in 3D and over time generate gigabytes of clinically relevant data per sample. There is a need for AI models to analyze this data and assist clinical workflow. The lack of…

Pulmonary pathologies are a significant global health concern, often leading to fatal outcomes if not diagnosed and treated promptly. Chest radiography serves as a primary diagnostic tool, but the availability of experienced radiologists…

Image and Video Processing · Electrical Eng. & Systems 2024-12-17 Abdelbaki Souid , Mohamed Hamroun , Soufiene Ben Othman , Hedi Sakli , Naceur Abdelkarim

The absence of adequately sufficient expert-level tumor annotations hinders the effectiveness of supervised learning based opportunistic cancer screening on medical imaging. Clinical reports (that are rich in descriptive textual details)…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Guangyu Guo , Jiawen Yao , Yingda Xia , Tony C. W. Mok , Zhilin Zheng , Junwei Han , Le Lu , Dingwen Zhang , Jian Zhou , Ling Zhang

Atypical mitotic figures (AMFs) are rare abnormal cell divisions associated with tumor aggressiveness and poor prognosis. Their detection remains a significant challenge due to subtle morphological cues, class imbalance, and inter-observer…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Lavish Ramchandani , Gunjan Deotale , Dev Kumar Das

In this study, we developed a deep-learning-based automatic detection algorithm (DLAD, Carebot AI CXR) to detect and localize seven specific radiological findings (atelectasis (ATE), consolidation (CON), pleural effusion (EFF), pulmonary…

Image and Video Processing · Electrical Eng. & Systems 2023-06-05 Daniel Kvak , Anna Chromcová , Petra Ovesná , Jakub Dandár , Marek Biroš , Robert Hrubý , Daniel Dufek , Marija Pajdaković

Foundation vision-language models (VLMs) excel on natural images, but their utility for biomedical microscopy remains underexplored. In this paper, we investigate how in-context learning enables state-of-the-art VLMs to perform few-shot…

Background: This study introduces a Vision-Language Model (VLM) leveraging SIGLIP and Gemma-3b architectures for automated acute tuberculosis (TB) screening. By integrating chest X-ray images and clinical notes, the model aims to enhance…

Reliable automated analysis of Optical Coherence Tomography (OCT) imaging is crucial for diagnosing retinal disorders but faces a critical barrier: the need for expensive, labor-intensive expert annotations. Supervised deep learning models…

Computer Vision and Pattern Recognition · Computer Science 2026-04-27 Tania Haghighi , Sina Gholami , Hamed Tabkhi , Minhaj Nur Alam

Oral mucosal diseases such as leukoplakia, oral lichen planus, and recurrent aphthous ulcers exhibit diverse and overlapping visual features, making diagnosis challenging for non-specialists. While vision-language models (VLMs) have shown…

Quantitative Methods · Quantitative Biology 2025-10-17 Jia Zhang , Bodong Du , Yitong Miao , Dongwei Sun , Xiangyong Cao

Deep learning models have achieved remarkable accuracy in chest X-ray diagnosis, yet their widespread clinical adoption remains limited by the black-box nature of their predictions. Clinicians require transparent, verifiable explanations to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Yiming Tang , Wenjia Zhong , Rushi Shah , Dianbo Liu

Medical image interpretation using deep learning has shown promise but often requires extensive expert-annotated datasets. To reduce this annotation burden, we develop an Image-Graph Contrastive Learning framework that pairs chest X-rays…

Image and Video Processing · Electrical Eng. & Systems 2024-05-17 Sameer Khanna , Daniel Michael , Marinka Zitnik , Pranav Rajpurkar

Mitotic figures are classified into typical and atypical variants, with atypical counts correlating strongly with tumor aggressiveness. Accurate differentiation is therefore essential for patient prognostication and resource allocation, yet…

Image and Video Processing · Electrical Eng. & Systems 2025-09-19 Mieko Ochi , Bae Yuan

Large Vision Language Models (LVLMs) show promise in medical applications, but their inability to faithfully ground responses in visual evidence raises serious concerns about clinical trustworthiness. While visual attribution methods are…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Guangzhi Xiong , Qiao Jin , Sanchit Sinha , Zhiyong Lu , Aidong Zhang

When analysing screening mammograms, radiologists can naturally process information across two ipsilateral views of each breast, namely the cranio-caudal (CC) and mediolateral-oblique (MLO) views. These multiple related images provide…

Computer Vision and Pattern Recognition · Computer Science 2022-09-22 Yuanhong Chen , Hu Wang , Chong Wang , Yu Tian , Fengbei Liu , Michael Elliott , Davis J. McCarthy , Helen Frazer , Gustavo Carneiro

Few-shot learning presents a critical solution for cancer diagnosis in computational pathology (CPath), addressing fundamental limitations in data availability, particularly the scarcity of expert annotations and patient privacy…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Zhengrui Guo , Conghao Xiong , Jiabo Ma , Qichen Sun , Lishuang Feng , Jinzhuo Wang , Hao Chen

The choice of input text prompt plays a critical role in the performance of Vision-Language Pretrained (VLP) models such as CLIP. We present APoLLo, a unified multi-modal approach that combines Adapter and Prompt learning for…

Machine Learning · Computer Science 2023-12-05 Sanjoy Chowdhury , Sayan Nag , Dinesh Manocha

Dermatological diagnosis represents a complex multimodal challenge that requires integrating visual features with specialized clinical knowledge. While vision-language pretraining (VLP) has advanced medical AI, its effectiveness in…

Computer Vision and Pattern Recognition · Computer Science 2025-05-15 Siyuan Yan , Xieji Li , Ming Hu , Yiwen Jiang , Zhen Yu , Zongyuan Ge

Medical image classification requires labeled, task-specific datasets which are used to train deep learning networks de novo, or to fine-tune foundation models. However, this process is computationally and technically demanding. In language…