中文
相关论文

相关论文: Continual Retinal Vision-Language Pre-training upo…

200 篇论文

Fine-tuning vision-language models (VLMs) such as CLIP often leads to catastrophic forgetting of pretrained knowledge. Prior work primarily aims to mitigate forgetting during adaptation; however, forgetting often remains inevitable during…

计算机视觉与模式识别 · 计算机科学 2026-02-27 Wenqing Wang , Da Li , Xiatian Zhu , Josef Kittler

Image fusion aims to combine information from multiple source images into a single one with more comprehensive informational content. Deep learning-based image fusion algorithms face significant challenges, including the lack of a…

计算机视觉与模式识别 · 计算机科学 2025-02-04 Haowen Bai , Zixiang Zhao , Jiangshe Zhang , Yichen Wu , Lilun Deng , Yukun Cui , Shuang Xu , Baisong Jiang

High myopia significantly increases the risk of irreversible vision loss. Traditional perimetry-based visual field (VF) assessment provides systematic quantification of visual loss but it is subjective and time-consuming. Consequently,…

计算机视觉与模式识别 · 计算机科学 2024-07-18 Zipei Yan , Zhile Liang , Zhengji Liu , Shuai Wang , Rachel Ka-Man Chun , Jizhou Li , Chea-su Kee , Dong Liang

Retinal imaging is fast, non-invasive, and widely available, offering quantifiable structural and vascular signals for ophthalmic and systemic health assessment. This accessibility creates an opportunity to study how quantitative retinal…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Zhonghua Wang , Lie Ju , Sijia Li , Wei Feng , Sijin Zhou , Ming Hu , Jianhao Xiong , Xiaoying Tang , Yifan Peng , Mingquan Lin , Yaodong Ding , Yong Zeng , Wenbin Wei , Li Dong , Zongyuan Ge

Prompt learning has become one of the most efficient paradigms for adapting large pre-trained vision-language models to downstream tasks. Current state-of-the-art methods, like CoOp and ProDA, tend to adopt soft prompts to learn an…

计算机视觉与模式识别 · 计算机科学 2023-03-31 Sifan Long , Zhen Zhao , Junkun Yuan , Zichang Tan , Jiangjiang Liu , Luping Zhou , Shengsheng Wang , Jingdong Wang

Large models have demonstrated exceptional generalization capabilities in computer vision and natural language processing. Recent efforts have focused on enhancing these models with multimodal processing abilities. However, addressing the…

计算机视觉与模式识别 · 计算机科学 2024-06-17 Hao Sun , Yu Song

Retinal vessel segmentation is a fundamental step in screening, diagnosis, and treatment of various cardiovascular and ophthalmic diseases. Robustness is one of the most critical requirements for practical utilization, since the test images…

图像与视频处理 · 电气工程与系统科学 2021-09-29 Xu Sun , Huihui Fang , Yehui Yang , Dongwei Zhu , Lei Wang , Junwei Liu , Yanwu Xu

Vision-language fine-tuning has emerged as an efficient paradigm for constructing multimodal foundation models. While textual context often highlights semantic relationships within an image, existing fine-tuning methods typically overlook…

计算机视觉与模式识别 · 计算机科学 2025-11-14 Xiangyang Wu , Liu Liu , Baosheng Yu , Jiayan Qiu , Zhenwei Shi

Retinal fundus photography offers a non-invasive way to diagnose and monitor a variety of retinal diseases, but is prone to inherent quality glitches arising from systemic imperfections or operator/patient-related factors. However,…

图像与视频处理 · 电气工程与系统科学 2024-09-13 Vamsi Krishna Vasa , Peijie Qiu , Wenhui Zhu , Yujian Xiong , Oana Dumitrascu , Yalin Wang

Multimodal learning, which involves integrating information from various modalities such as text, images, audio, and video, is pivotal for numerous complex tasks like visual question answering, cross-modal retrieval, and caption generation.…

计算机视觉与模式识别 · 计算机科学 2025-07-29 G. Thomas Hudson , Dean Slack , Thomas Winterbottom , Jamie Sterling , Chenghao Xiao , Junjie Shentu , Noura Al Moubayed

Retinal fundus images are widely used for the clinical screening and diagnosis of eye diseases. However, fundus images captured by operators with various levels of experience have a large variation in quality. Low-quality fundus images…

图像与视频处理 · 电气工程与系统科学 2020-12-10 Ziyi Shen , Huazhu Fu , Jianbing Shen , Ling Shao

Foundation models are used to extract transferable representations from large amounts of unlabeled data, typically via self-supervised learning (SSL). However, many of these models rely on architectures that offer limited interpretability,…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Samuel Ofosu Mensah , Camila Roa , Kerol Djoumessi , Philipp Berens

With the recent progress in large-scale vision and language representation learning, Vision Language Pre-training (VLP) models have achieved promising improvements on various multi-modal downstream tasks. Albeit powerful, these models have…

计算机视觉与模式识别 · 计算机科学 2023-08-08 Jiahua Rao , Zifei Shan , Longpo Liu , Yao Zhou , Yuedong Yang

Retinal fundus photography enhancement is important for diagnosing and monitoring retinal diseases. However, early approaches to retinal image enhancement, such as those based on Generative Adversarial Networks (GANs), often struggle to…

图像与视频处理 · 电气工程与系统科学 2024-11-05 Xuanzhao Dong , Wenhui Zhu , Xin Li , Guoxin Sun , Yi Su , Oana M. Dumitrascu , Yalin Wang

Foundation models encompass an extensive knowledge base and offer remarkable transferability. However, this knowledge becomes outdated or insufficient over time. The challenge lies in continuously updating foundation models to accommodate…

计算机视觉与模式识别 · 计算机科学 2024-04-22 Wenxuan Zhang , Paul Janson , Rahaf Aljundi , Mohamed Elhoseiny

Large-scale contrastive pre-training produces powerful Vision-and-Language Models (VLMs) capable of generating representations (embeddings) effective for a wide variety of visual and multimodal tasks. However, these pretrained embeddings…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Nikolaos-Antonios Ypsilantis , Kaifeng Chen , André Araujo , Ondřej Chum

Contrastive learning based vision-language joint pre-training has emerged as a successful representation learning strategy. In this paper, we present a prototype representation learning framework incorporating both global and local…

计算机视觉与模式识别 · 计算机科学 2024-03-12 Pujin Cheng , Li Lin , Junyan Lyu , Yijin Huang , Wenhan Luo , Xiaoying Tang

Recent advancements in ophthalmology foundation models such as RetFound have demonstrated remarkable diagnostic capabilities but require massive datasets for effective pre-training, creating significant barriers for development and…

图像与视频处理 · 电气工程与系统科学 2025-03-25 Qingshan Hou , Meng Wang , Peng Cao , Zou Ke , Xiaoli Liu , Huazhu Fu , Osmar R. Zaiane

The advent of vision-language pre-training techniques enhanced substantial progress in the development of models for image captioning. However, these models frequently produce generic captions and may omit semantically important image…

计算机视觉与模式识别 · 计算机科学 2023-11-17 Noam Rotstein , David Bensaid , Shaked Brody , Roy Ganz , Ron Kimmel

Real-world non-mydriatic retinal fundus photography is prone to artifacts, imperfections and low-quality when certain ocular or systemic co-morbidities exist. Artifacts may result in inaccuracy or ambiguity in clinical diagnoses. In this…

图像与视频处理 · 电气工程与系统科学 2023-02-07 Wenhui Zhu , Peijie Qiu , Mohammad Farazi , Keshav Nandakumar , Oana M. Dumitrascu , Yalin Wang