English
Related papers

Related papers: SimCroP: Radiograph Representation Learning with S…

200 papers

While Vision-Language Models (VLMs) have achieved notable progress in computational pathology (CPath), the gigapixel scale and spatial heterogeneity of Whole Slide Images (WSIs) continue to pose challenges for multimodal understanding.…

Computer Vision and Pattern Recognition · Computer Science 2025-12-22 Fengchun Liu , Songhan Jiang , Linghan Cai , Ziyue Wang , Yongbing Zhang

Vision Language Foundation Models based on CLIP architecture for remote sensing primarily rely on short text captions, which often result in incomplete semantic representations. Although longer captions convey richer information, existing…

Computer Vision and Pattern Recognition · Computer Science 2025-10-30 Weizhi Chen , Yupeng Deng , Jin Wei , Jingbo Chen , Jiansheng Chen , Yuman Feng , Zhihao Xi , Diyou Liu , Kai Li , Yu Meng

The language-guided robot grasping task requires a robot agent to integrate multimodal information from both visual and linguistic inputs to predict actions for target-driven grasping. While recent approaches utilizing Multimodal Large…

Robotics · Computer Science 2025-02-10 Houjian Yu , Mingen Li , Alireza Rezazadeh , Yang Yang , Changhyun Choi

As medical imaging is central to diagnostic processes, automating the generation of radiology reports has become increasingly relevant to assist radiologists with their heavy workloads. Most current methods rely solely on global image…

Computer Vision and Pattern Recognition · Computer Science 2025-08-08 Hamza Kalisch , Fabian Hörst , Jens Kleesiek , Ken Herrmann , Constantin Seibold

Vision Foundation Model (VFM) such as the Segment Anything Model (SAM) and Contrastive Language-Image Pre-training Model (CLIP) has shown promising performance for segmentation and detection tasks. However, although SAM excels in…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Kunliang Liu , Jianming Wang , Rize Jin , Wonjun Hwang , Tae-Sun Chung

Medical image segmentation remains challenging due to limited annotations for training, ambiguous anatomical features, and domain shifts. While vision-language models such as CLIP offer strong cross-modal representations, their potential…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Taha Koleilat , Hojat Asgariandehkordi , Omid Nejati Manzari , Berardino Barile , Yiming Xiao , Hassan Rivaz

Recent advances in prototypical learning have shown remarkable potential to provide useful decision interpretations associating activation maps and predictions with class-specific training prototypes. Such prototypical learning has been…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Chong Wang , Fengbei Liu , Yuanhong Chen , Helen Frazer , Gustavo Carneiro

Automatic summarization of radiology reports is an essential application to reduce the burden on physicians. Previous studies have widely used the "pre-training, fine-tuning" strategy to adapt large language models (LLMs) for summarization.…

Computation and Language · Computer Science 2026-04-13 Mengxian Lyu , Cheng Peng , Ziyi Chen , Mengyuan Zhang , Jieting Li Lu , Yonghui Wu

Vision-language pretraining models have made significant progress in bridging remote sensing imagery with natural language. However, existing approaches often fail to effectively integrate multi-granular visual and textual information,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Xiao Yang , Ronghao Fu , Zhuoran Duan , Zhiwen Lin , Xueyan Liu , Bo Yang

Medical image segmentation driven by free-text clinical instructions is a critical frontier in computer-aided diagnosis. However, existing multimodal and foundation models struggle with the semantic ambiguity of clinical reports and fail to…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Chenyu Xue , Yiran Liu , Mian Zhou , Jionglong Su , Zhixiang Lu

Radiology reports are detailed text descriptions of the content of medical scans. Each report describes the presence/absence and location of relevant clinical findings, commonly including comparison with prior exams of the same patient to…

Computer Vision and Pattern Recognition · Computer Science 2023-10-10 Francesco Dalla Serra , Chaoyang Wang , Fani Deligianni , Jeffrey Dalton , Alison Q O'Neil

Learning medical visual representations from paired images and reports is a promising direction in representation learning. However, current vision-language pretraining methods in the medical domain often simplify clinical reports into…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 Wei Li , Xun Gong , Jiao Li , Xiaobin Sun

Non-contrast CT (NCCT) imaging may reduce image contrast and anatomical visibility, potentially increasing diagnostic uncertainty. In contrast, contrast-enhanced CT (CECT) facilitates the observation of regions of interest (ROI). Leading…

Image and Video Processing · Electrical Eng. & Systems 2024-11-18 Tingyi Lin , Pengju Lyu , Jie Zhang , Yuqing Wang , Cheng Wang , Jianjun Zhu

Structural magnetic resonance imaging (sMRI) provides accurate estimates of the brain's structural organization and learning invariant brain representations from sMRI is an enduring issue in neuroscience. Previous deep representation…

Computer Vision and Pattern Recognition · Computer Science 2023-06-21 Ning Jiang , Gongshu Wang , Tianyi Yan

Vision-language pre-training (VLP) has great potential for developing multifunctional and general medical diagnostic capabilities. However, aligning medical images with a low signal-to-noise ratio (SNR) to reports with a high SNR presents a…

Image and Video Processing · Electrical Eng. & Systems 2025-08-07 Weiwei Cao , Jianpeng Zhang , Zhongyi Shui , Sinuo Wang , Zeli Chen , Xi Li , Le Lu , Xianghua Ye , Tingbo Liang , Qi Zhang , Ling Zhang

Recent advances in general-purpose foundation models have stimulated the development of large biological sequence models. While natural language shows symbolic granularity (characters, words, sentences), biological sequences exhibit…

Machine Learning · Computer Science 2026-03-24 Hanlin Xiao , Rainer Breitling , Eriko Takano , Mauricio A. Álvarez

Pre-trained models, e.g., from ImageNet, have proven to be effective in boosting the performance of many downstream applications. It is too demanding to acquire large-scale annotations to build such models for medical imaging. Meanwhile,…

Computer Vision and Pattern Recognition · Computer Science 2021-03-31 Xiaosong Wang , Ziyue Xu , Leo Tam , Dong Yang , Daguang Xu

Recent years have witnessed the fast development of large-scale pre-training frameworks that can extract multi-modal representations in a unified form and achieve promising performances when transferred to downstream tasks. Nevertheless,…

Computer Vision and Pattern Recognition · Computer Science 2022-10-18 Xuran Pan , Tianzhu Ye , Dongchen Han , Shiji Song , Gao Huang

Medical report generation is a challenging task since it is time-consuming and requires expertise from experienced radiologists. The goal of medical report generation is to accurately capture and describe the image findings. Previous works…

Computer Vision and Pattern Recognition · Computer Science 2023-01-10 Yu-Jen Chen , Wei-Hsiang Shen , Hao-Wei Chung , Ching-Hao Chiu , Da-Cheng Juan , Tsung-Ying Ho , Chi-Tung Cheng , Meng-Lin Li , Tsung-Yi Ho

Medical vision-language pretraining (VLP) that leverages naturally-paired medical image-report data is crucial for medical image analysis. However, existing methods struggle to accurately characterize associations between images and…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Xinjie Liang , Xiangyu Li , Fanding Li , Jie Jiang , Qing Dong , Wei Wang , Kuanquan Wang , Suyu Dong , Gongning Luo , Shuo Li
‹ Prev 1 4 5 6 7 8 10 Next ›