English
Related papers

Related papers: MeDSLIP: Medical Dual-Stream Language-Image Pre-tr…

200 papers

Vision Language Models (VLMs) like CLIP have attracted substantial attention in pathology, serving as backbones for applications such as zero-shot image classification and Whole Slide Image (WSI) analysis. Additionally, they can function as…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Yuxuan Sun , Yunlong Zhang , Yixuan Si , Chenglu Zhu , Zhongyi Shui , Kai Zhang , Jingxiong Li , Xingheng Lyu , Tao Lin , Lin Yang

The high incidence and mortality rates associated with respiratory diseases underscores the importance of early screening. Machine learning models can automate clinical consultations and auscultation, offering vital support in this area.…

Machine Learning · Computer Science 2024-10-10 Yuwei Zhang , Tong Xia , Aaqib Saeed , Cecilia Mascolo

The scarcity of annotated data has sparked significant interest in unsupervised pre-training methods that leverage medical reports as auxiliary signals for medical visual representation learning. However, existing research overlooks the…

Computer Vision and Pattern Recognition · Computer Science 2024-02-06 Zhe Li , Laurence T. Yang , Bocheng Ren , Xin Nie , Zhangyang Gao , Cheng Tan , Stan Z. Li

Medical image segmentation typically relies solely on visual data, overlooking the rich textual information clinicians use for diagnosis. Vision-language models attempt to bridge this gap, but existing approaches often process visual and…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Rafi Ibn Sultan , Hui Zhu , Chengyin Li , Dongxiao Zhu

Vision-language foundation models (VLMs) have shown great potential in feature transfer and generalization across a wide spectrum of medical-related downstream tasks. However, fine-tuning these models is resource-intensive due to their…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Ye Du , Nanxi Yu , Shujun Wang

Vision-language modeling (VLM) aims to bridge the information gap between images and natural language. Under the new paradigm of first pre-training on massive image-text pairs and then fine-tuning on task-specific data, VLM in the remote…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Xingxing Weng , Chao Pang , Gui-Song Xia

The recent rapid advancements in language models (LMs) have garnered attention in medical time series-text multimodal learning. However, existing contrastive learning-based and prompt-based LM approaches tend to be biased, often assigning a…

Machine Learning · Computer Science 2025-09-09 Jiexia Ye , Weiqi Zhang , Ziyue Li , Jia Li , Meng Zhao , Fugee Tsung

One of the biggest challenges that prohibit the use of many current NLP methods in clinical settings is the availability of public datasets. In this work, we present MeDAL, a large medical text dataset curated for abbreviation…

Computation and Language · Computer Science 2020-12-29 Zhi Wen , Xing Han Lu , Siva Reddy

Research on Multi-modal Large Language Models (MLLMs) towards the multi-image cross-modal instruction has received increasing attention and made significant progress, particularly in scenarios involving closely resembling images (e.g.,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-26 Tao Wu , Mengze Li , Jingyuan Chen , Wei Ji , Wang Lin , Jinyang Gao , Kun Kuang , Zhou Zhao , Fei Wu

Medical Vision-Language Pre-training (MedVLP) has made significant progress in enabling zero-shot tasks for medical image understanding. However, training MedVLP models typically requires large-scale datasets with paired, high-quality…

Computer Vision and Pattern Recognition · Computer Science 2025-02-26 Che Liu , Zhongwei Wan , Haozhe Wang , Yinda Chen , Talha Qaiser , Chen Jin , Fariba Yousefi , Nikolay Burlutskiy , Rossella Arcucci

Medical image analysis is essential in modern healthcare. Deep learning has redirected research focus toward complex medical multimodal tasks, including report generation and visual question answering. Traditional task-specific models often…

Computer Vision and Pattern Recognition · Computer Science 2025-09-15 Yiming Shi , Shaoshuai Yang , Xun Zhu , Haoyu Wang , Xiangling Fu , Miao Li , Ji Wu

Accurate diagnosis of skin diseases remains a significant challenge due to the complex and diverse visual features present in dermatoscopic images, often compounded by a lack of interpretability in existing purely visual diagnostic models.…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Kexin Yu , Zihan Xu , Jialei Xie , Carter Adams

Semi-supervised medical image segmentation (SSMIS) uses consistency learning to regularize model training, which alleviates the burden of pixel-wise manual annotations. However, it often suffers from error supervision from low-quality…

Computer Vision and Pattern Recognition · Computer Science 2024-12-18 Qingtao Pan , Wenhao Qiao , Jingjiao Lou , Bing Ji , Shuo Li

The inability to interpret the model prediction in semantically and visually meaningful ways is a well-known shortcoming of most existing computer-aided diagnosis methods. In this paper, we propose MDNet to establish a direct multimodal…

Computer Vision and Pattern Recognition · Computer Science 2017-07-11 Zizhao Zhang , Yuanpu Xie , Fuyong Xing , Mason McGough , Lin Yang

Medical Vision-Language Pretraining (MedVLP) shows promise in learning generalizable and transferable visual representations from paired and unpaired medical images and reports. MedVLP can provide useful features to downstream tasks and…

Computer Vision and Pattern Recognition · Computer Science 2024-10-30 Yang Zhou , Tan Li Hui Faith , Yanyu Xu , Sicong Leng , Xinxing Xu , Yong Liu , Rick Siow Mong Goh

Conversational artificial intelligence has the potential to assist users in preliminary medical consultations, particularly in settings where access to healthcare professionals is limited. However, many existing medical dialogue systems…

Computation and Language · Computer Science 2026-03-26 Shubham Kumar Nigam , Suparnojit Sarkar , Piyush Patel

The complexity and heterogeneity of data in many real-world applications pose significant challenges for traditional machine learning and signal processing techniques. For instance, in medicine, effective analysis of diverse physiological…

Machine Learning · Computer Science 2024-08-16 Nimeesha Chan , Felix Parker , William Bennett , Tianyi Wu , Mung Yao Jia , James Fackler , Kimia Ghobadi

Multi-modal large language models (MLLMs) have shown promise in advancing healthcare. However, most existing models remain confined to single-image understanding, which greatly limits their applicability in clinical workflows. In practice,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Zhen Chen , Yihang Fu , Gabriel Madera , Mauro Giuffre , Serina Applebaum , Hyunjae Kim , Hua Xu , Qingyu Chen

Whole Slide Images (WSIs) exhibit hierarchical structure, where diagnostic information emerges from cellular morphology, regional tissue organization, and global context. Existing Computational Pathology (CPath) Multimodal Large Language…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Basit Alawode , Arif Mahmood , Muaz Khalifa Al-Radi , Shahad Albastaki , Asim Khan , Muhammad Bilal , Moshira Ali Abdalla , Mohammed Bennamoun , Sajid Javed

Renal transplantation emerges as the most effective solution for end-stage renal disease. Occurring from complex causes, a substantial risk of transplant chronic dysfunction persists and may lead to graft loss. Medical imaging plays a…

Computer Vision and Pattern Recognition · Computer Science 2023-05-02 Leo Milecki , Vicky Kalogeiton , Sylvain Bodard , Dany Anglicheau , Jean-Michel Correas , Marc-Olivier Timsit , Maria Vakalopoulou