English
Related papers

Related papers: Continual Retinal Vision-Language Pre-training upo…

200 papers

Fine-tuning vision-language models (VLMs) such as CLIP often leads to catastrophic forgetting of pretrained knowledge. Prior work primarily aims to mitigate forgetting during adaptation; however, forgetting often remains inevitable during…

Computer Vision and Pattern Recognition · Computer Science 2026-02-27 Wenqing Wang , Da Li , Xiatian Zhu , Josef Kittler

Image fusion aims to combine information from multiple source images into a single one with more comprehensive informational content. Deep learning-based image fusion algorithms face significant challenges, including the lack of a…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Haowen Bai , Zixiang Zhao , Jiangshe Zhang , Yichen Wu , Lilun Deng , Yukun Cui , Shuang Xu , Baisong Jiang

High myopia significantly increases the risk of irreversible vision loss. Traditional perimetry-based visual field (VF) assessment provides systematic quantification of visual loss but it is subjective and time-consuming. Consequently,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-18 Zipei Yan , Zhile Liang , Zhengji Liu , Shuai Wang , Rachel Ka-Man Chun , Jizhou Li , Chea-su Kee , Dong Liang

Retinal imaging is fast, non-invasive, and widely available, offering quantifiable structural and vascular signals for ophthalmic and systemic health assessment. This accessibility creates an opportunity to study how quantitative retinal…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Zhonghua Wang , Lie Ju , Sijia Li , Wei Feng , Sijin Zhou , Ming Hu , Jianhao Xiong , Xiaoying Tang , Yifan Peng , Mingquan Lin , Yaodong Ding , Yong Zeng , Wenbin Wei , Li Dong , Zongyuan Ge

Prompt learning has become one of the most efficient paradigms for adapting large pre-trained vision-language models to downstream tasks. Current state-of-the-art methods, like CoOp and ProDA, tend to adopt soft prompts to learn an…

Computer Vision and Pattern Recognition · Computer Science 2023-03-31 Sifan Long , Zhen Zhao , Junkun Yuan , Zichang Tan , Jiangjiang Liu , Luping Zhou , Shengsheng Wang , Jingdong Wang

Large models have demonstrated exceptional generalization capabilities in computer vision and natural language processing. Recent efforts have focused on enhancing these models with multimodal processing abilities. However, addressing the…

Computer Vision and Pattern Recognition · Computer Science 2024-06-17 Hao Sun , Yu Song

Retinal vessel segmentation is a fundamental step in screening, diagnosis, and treatment of various cardiovascular and ophthalmic diseases. Robustness is one of the most critical requirements for practical utilization, since the test images…

Image and Video Processing · Electrical Eng. & Systems 2021-09-29 Xu Sun , Huihui Fang , Yehui Yang , Dongwei Zhu , Lei Wang , Junwei Liu , Yanwu Xu

Vision-language fine-tuning has emerged as an efficient paradigm for constructing multimodal foundation models. While textual context often highlights semantic relationships within an image, existing fine-tuning methods typically overlook…

Computer Vision and Pattern Recognition · Computer Science 2025-11-14 Xiangyang Wu , Liu Liu , Baosheng Yu , Jiayan Qiu , Zhenwei Shi

Retinal fundus photography offers a non-invasive way to diagnose and monitor a variety of retinal diseases, but is prone to inherent quality glitches arising from systemic imperfections or operator/patient-related factors. However,…

Image and Video Processing · Electrical Eng. & Systems 2024-09-13 Vamsi Krishna Vasa , Peijie Qiu , Wenhui Zhu , Yujian Xiong , Oana Dumitrascu , Yalin Wang

Multimodal learning, which involves integrating information from various modalities such as text, images, audio, and video, is pivotal for numerous complex tasks like visual question answering, cross-modal retrieval, and caption generation.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 G. Thomas Hudson , Dean Slack , Thomas Winterbottom , Jamie Sterling , Chenghao Xiao , Junjie Shentu , Noura Al Moubayed

Retinal fundus images are widely used for the clinical screening and diagnosis of eye diseases. However, fundus images captured by operators with various levels of experience have a large variation in quality. Low-quality fundus images…

Image and Video Processing · Electrical Eng. & Systems 2020-12-10 Ziyi Shen , Huazhu Fu , Jianbing Shen , Ling Shao

Foundation models are used to extract transferable representations from large amounts of unlabeled data, typically via self-supervised learning (SSL). However, many of these models rely on architectures that offer limited interpretability,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Samuel Ofosu Mensah , Camila Roa , Kerol Djoumessi , Philipp Berens

With the recent progress in large-scale vision and language representation learning, Vision Language Pre-training (VLP) models have achieved promising improvements on various multi-modal downstream tasks. Albeit powerful, these models have…

Computer Vision and Pattern Recognition · Computer Science 2023-08-08 Jiahua Rao , Zifei Shan , Longpo Liu , Yao Zhou , Yuedong Yang

Retinal fundus photography enhancement is important for diagnosing and monitoring retinal diseases. However, early approaches to retinal image enhancement, such as those based on Generative Adversarial Networks (GANs), often struggle to…

Image and Video Processing · Electrical Eng. & Systems 2024-11-05 Xuanzhao Dong , Wenhui Zhu , Xin Li , Guoxin Sun , Yi Su , Oana M. Dumitrascu , Yalin Wang

Foundation models encompass an extensive knowledge base and offer remarkable transferability. However, this knowledge becomes outdated or insufficient over time. The challenge lies in continuously updating foundation models to accommodate…

Computer Vision and Pattern Recognition · Computer Science 2024-04-22 Wenxuan Zhang , Paul Janson , Rahaf Aljundi , Mohamed Elhoseiny

Large-scale contrastive pre-training produces powerful Vision-and-Language Models (VLMs) capable of generating representations (embeddings) effective for a wide variety of visual and multimodal tasks. However, these pretrained embeddings…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Nikolaos-Antonios Ypsilantis , Kaifeng Chen , André Araujo , Ondřej Chum

Contrastive learning based vision-language joint pre-training has emerged as a successful representation learning strategy. In this paper, we present a prototype representation learning framework incorporating both global and local…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Pujin Cheng , Li Lin , Junyan Lyu , Yijin Huang , Wenhan Luo , Xiaoying Tang

Recent advancements in ophthalmology foundation models such as RetFound have demonstrated remarkable diagnostic capabilities but require massive datasets for effective pre-training, creating significant barriers for development and…

Image and Video Processing · Electrical Eng. & Systems 2025-03-25 Qingshan Hou , Meng Wang , Peng Cao , Zou Ke , Xiaoli Liu , Huazhu Fu , Osmar R. Zaiane

The advent of vision-language pre-training techniques enhanced substantial progress in the development of models for image captioning. However, these models frequently produce generic captions and may omit semantically important image…

Computer Vision and Pattern Recognition · Computer Science 2023-11-17 Noam Rotstein , David Bensaid , Shaked Brody , Roy Ganz , Ron Kimmel

Real-world non-mydriatic retinal fundus photography is prone to artifacts, imperfections and low-quality when certain ocular or systemic co-morbidities exist. Artifacts may result in inaccuracy or ambiguity in clinical diagnoses. In this…

Image and Video Processing · Electrical Eng. & Systems 2023-02-07 Wenhui Zhu , Peijie Qiu , Mohammad Farazi , Keshav Nandakumar , Oana M. Dumitrascu , Yalin Wang