English
Related papers

Related papers: Knowledge-enhanced Visual-Language Pretraining for…

200 papers

We develop an approach to learning visual representations that embraces multimodal data, driven by a combination of intra- and inter-modal similarity preservation objectives. Unlike existing visual pre-training methods, which solve a proxy…

Computer Vision and Pattern Recognition · Computer Science 2021-04-28 Xin Yuan , Zhe Lin , Jason Kuen , Jianming Zhang , Yilin Wang , Michael Maire , Ajinkya Kale , Baldo Faieta

Deep learning models can be applied successfully in real-work problems; however, training most of these models requires massive data. Recent methods use language and vision, but unfortunately, they rely on datasets that are not usually…

Computer Vision and Pattern Recognition · Computer Science 2023-01-27 Nathan Hadjiyski , Ali Vosoughi , Axel Wismueller

Despite deep convolutional neural networks boost the performance of image classification and segmentation in digital pathology analysis, they are usually weak in interpretability for clinical applications or require heavy annotations to…

Computer Vision and Pattern Recognition · Computer Science 2020-02-10 Yongxiang Huang , Albert C. S. Chung

Accurate disease interpretation from radiology remains challenging due to imaging heterogeneity. Achieving expert-level diagnostic decisions requires integration of subtle image features with clinical knowledge. Yet major vision-language…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Difei Gu , Yunhe Gao , Mu Zhou , Dimitris Metaxas

Computational pathology (CPath) digitizes pathology slides into whole slide images (WSIs), enabling analysis for critical healthcare tasks such as cancer diagnosis and prognosis. However, WSIs possess extremely long sequence lengths (up to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Wenhao Tang , Heng Fang , Ge Wu , Xiang Li , Ming-Ming Cheng

Pathology text mining is a challenging task given the reporting variability and constant new findings in cancer sub-type definitions. However, successful text mining of a large pathology database can play a critical role to advance 'big…

Computation and Language · Computer Science 2022-05-17 Thiago Santos , Amara Tariq , Susmita Das , Kavyasree Vayalpati , Geoffrey H. Smith , Hari Trivedi , Imon Banerjee

This paper explores training medical vision-language models (VLMs) -- where the visual and language inputs are embedded into a common space -- with a particular focus on scenarios where training data is limited, as is often the case in…

Computer Vision and Pattern Recognition · Computer Science 2023-04-03 Rhydian Windsor , Amir Jamaludin , Timor Kadir , Andrew Zisserman

Though deep learning has shown successful performance in classifying the label and severity stage of certain disease, most of them give few evidence on how to make prediction. Here, we propose to exploit the interpretability of deep…

Computer Vision and Pattern Recognition · Computer Science 2020-03-16 Yuhao Niu , Lin Gu , Feng Lu , Feifan Lv , Zongji Wang , Imari Sato , Zijian Zhang , Yangyan Xiao , Xunzhang Dai , Tingting Cheng

Computational pathology and whole-slide image (WSI) analysis are pivotal in cancer diagnosis and prognosis. However, the ultra-high resolution of WSIs presents significant modeling challenges. Recent advancements in pathology foundation…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Honglin Li , Zhongyi Shui , Yunlong Zhang , Chenglu Zhu , Lin Yang

Medical Vision-Language Pre-training (VLP) learns representations jointly from medical images and paired radiology reports. It typically requires large-scale paired image-text datasets to achieve effective pre-training for both the image…

Computer Vision and Pattern Recognition · Computer Science 2024-05-01 Che Liu , Anand Shah , Wenjia Bai , Rossella Arcucci

Pathology is experiencing rapid digital transformation driven by whole-slide imaging and artificial intelligence (AI). While deep learning-based computational pathology has achieved notable success, traditional models primarily focus on…

Representation learning from Gigapixel Whole Slide Images (WSI) poses a significant challenge in computational pathology due to the complicated nature of tissue structures and the scarcity of labeled data. Multi-instance learning methods…

Image and Video Processing · Electrical Eng. & Systems 2024-05-28 Ali Nasiri-Sarvi , Vincent Quoc-Huy Trinh , Hassan Rivaz , Mahdi S. Hosseini

Brain CT report generation is significant to aid physicians in diagnosing cranial diseases. Recent studies concentrate on handling the consistency between visual and textual pathological features to improve the coherence of report. However,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-02 Chengxin Zheng , Junzhong Ji , Yanzhao Shi , Xiaodan Zhang , Liangqiong Qu

Learning visual representations of medical images (e.g., X-rays) is core to medical image understanding but its progress has been held back by the scarcity of human annotations. Existing work commonly relies on fine-tuning weights…

Computer Vision and Pattern Recognition · Computer Science 2022-09-21 Yuhao Zhang , Hang Jiang , Yasuhide Miura , Christopher D. Manning , Curtis P. Langlotz

Recently, histopathology vision-language foundation models (VLMs) have gained popularity due to their enhanced performance and generalizability across different downstream tasks. However, most existing histopathology benchmarks are either…

Image and Video Processing · Electrical Eng. & Systems 2025-03-18 Roba Al Majzoub , Hashmat Malik , Muzammal Naseer , Zaigham Zaheer , Tariq Mahmood , Salman Khan , Fahad Khan

Due to the increasing workload of pathologists, the need for automation to support diagnostic tasks and quantitative biomarker evaluation is becoming more and more apparent. Foundation models have the potential to improve generalizability…

Image and Video Processing · Electrical Eng. & Systems 2025-01-13 Till Nicke , Jan Raphael Schaefer , Henning Hoefener , Friedrich Feuerhake , Dorit Merhof , Fabian Kiessling , Johannes Lotz

Medical image classification plays a crucial role in clinical decision-making, yet most models are constrained to a fixed set of predefined classes, limiting their adaptability to new conditions. Contrastive Language-Image Pretraining…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Stefan Denner , Markus Bujotzek , Dimitrios Bounias , David Zimmerer , Raphael Stock , Klaus Maier-Hein

With the increase in the use of deep learning for computer-aided diagnosis in medical images, the criticism of the black-box nature of the deep learning models is also on the rise. The medical community needs interpretable models for both…

Image and Video Processing · Electrical Eng. & Systems 2020-12-21 Mookund Sureka , Abhijeet Patil , Deepak Anand , Amit Sethi

Self-supervised learning methods for computer vision have demonstrated the effectiveness of pre-training feature representations, resulting in well-generalizing Deep Neural Networks, even if the annotated data are limited. However,…

Computer Vision and Pattern Recognition · Computer Science 2021-08-25 Dmitrii Shubin , Danny Eytan , Sebastian D. Goodfellow

Vision-language models (VLMs) embed aligned image-text pairs into a joint space but often rely on deterministic embeddings, assuming a one-to-one correspondence between images and texts. This oversimplifies real-world relationships, which…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Sanghyuk Chun , Wonjae Kim , Song Park , Sangdoo Yun