English
Related papers

Related papers: UrFound: Towards Universal Retinal Foundation Mode…

200 papers

Wireless foundation models (WFMs) have recently demonstrated promising capabilities, jointly performing multiple wireless functions and adapting effectively to new environments. However, while current WFMs process only one modality,…

Signal Processing · Electrical Eng. & Systems 2026-02-20 Ahmed Aboulfotouh , Hatem Abou-Zeid

Image inpainting is an ill-posed problem to recover missing or damaged image content based on incomplete images with masks. Previous works usually predict the auxiliary structures (e.g., edges, segmentation and contours) to help fill…

Computer Vision and Pattern Recognition · Computer Science 2022-08-26 Yongsheng Yu , Dawei Du , Libo Zhang , Tiejian Luo

Despite rapid advances in face recognition, there remains a clear gap between the performance of still image-based face recognition and video-based face recognition, due to the vast difference in visual quality between the domains and the…

Computer Vision and Pattern Recognition · Computer Science 2017-08-15 Kihyuk Sohn , Sifei Liu , Guangyu Zhong , Xiang Yu , Ming-Hsuan Yang , Manmohan Chandraker

Visible-infrared image fusion is crucial in key applications such as autonomous driving and nighttime surveillance. Its main goal is to integrate multimodal information to produce enhanced images that are better suited for downstream tasks.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Xiaopeng Liu , Yupei Lin , Sen Zhang , Xiao Wang , Yukai Shi , Liang Lin

Retinopathy represents a group of retinal diseases that, if not treated timely, can cause severe visual impairments or even blindness. Many researchers have developed autonomous systems to recognize retinopathy via fundus and optical…

Image and Video Processing · Electrical Eng. & Systems 2021-11-05 Taimur Hassan , Bilal Hassan , Muhammad Usman Akram , Shahrukh Hashmi , Abdel Hakim Taguri , Naoufel Werghi

Multimodal foundation models offer promising advancements for enhancing driving perception systems, but their high computational and financial costs pose challenges. We develop a method that leverages foundation models to refine predictions…

Computer Vision and Pattern Recognition · Computer Science 2024-10-03 Yunhao Yang , Yuxin Hu , Mao Ye , Zaiwei Zhang , Zhichao Lu , Yi Xu , Ufuk Topcu , Ben Snyder

This paper proposes a self-supervised approach to learn universal facial representations from videos, that can transfer across a variety of facial analysis tasks such as Facial Attribute Recognition (FAR), Facial Expression Recognition…

Computer Vision and Pattern Recognition · Computer Science 2023-03-23 Zhixi Cai , Shreya Ghosh , Kalin Stefanov , Abhinav Dhall , Jianfei Cai , Hamid Rezatofighi , Reza Haffari , Munawar Hayat

Recently, text-to-image denoising diffusion probabilistic models (DDPMs) have demonstrated impressive image generation capabilities and have also been successfully applied to image inpainting. However, in practice, users often require more…

Computer Vision and Pattern Recognition · Computer Science 2023-10-12 Shiyuan Yang , Xiaodong Chen , Jing Liao

Semantic analysis on visible (RGB) and infrared (IR) images has gained significant attention due to their enhanced accuracy and robustness under challenging conditions including low-illumination and adverse weather. However, due to the lack…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Maoxun Yuan , Bo Cui , Tianyi Zhao , Jiayi Wang , Shan Fu , Xue Yang , Xingxing Wei

Tabular Foundation Models have recently established the state of the art in supervised tabular learning, by leveraging pretraining to learn generalizable representations of numerical and categorical structured data. However, they lack…

Multimodal foundation models offer a promising framework for robotic perception and planning by processing sensory inputs to generate actionable plans. However, addressing uncertainty in both perception (sensory interpretation) and…

Robotics · Computer Science 2025-04-18 Neel P. Bhatt , Yunhao Yang , Rohan Siva , Daniel Milan , Ufuk Topcu , Zhangyang Wang

Many real-world visual recognition use-cases can not directly benefit from state-of-the-art CNN-based approaches because of the lack of many annotated data. The usual approach to deal with this is to transfer a representation pre-learned on…

Computer Vision and Pattern Recognition · Computer Science 2018-10-05 Julien Girard , Youssef Tamaazousti , Hervé Le Borgne , Céline Hudelot

Deep learning is currently the state-of-the-art for automated detection of referable diabetic retinopathy (DR) from color fundus photographs (CFP). While the general interest is put on improving results through methodological innovations,…

Image and Video Processing · Electrical Eng. & Systems 2022-10-10 Tomás Castilla , Marcela S. Martínez , Mercedes Leguía , Ignacio Larrabide , José Ignacio Orlando

In the field of face recognition, a model learns to distinguish millions of face images with fewer dimensional embedding features, and such vast information may not be properly encoded in the conventional model with a single branch. We…

Computer Vision and Pattern Recognition · Computer Science 2020-05-26 Yonghyun Kim , Wonpyo Park , Myung-Cheol Roh , Jongju Shin

Recent advances in multimodal large language models (MLLMs) have demonstrated strong capabilities in understanding general visual content. However, these general-domain MLLMs perform poorly in face perception tasks, often producing…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Jingzhi Li , Changjiang Luo , Ruoyu Chen , Hua Zhang , Wenqi Ren , Jianhou Gan , Xiaochun Cao

Current machine learning models for vision are often highly specialized and limited to a single modality and task. In contrast, recent large language models exhibit a wide range of capabilities, hinting at a possibility for similarly…

Computer Vision and Pattern Recognition · Computer Science 2023-12-12 David Mizrahi , Roman Bachmann , Oğuzhan Fatih Kar , Teresa Yeo , Mingfei Gao , Afshin Dehghan , Amir Zamir

Multimodal foundation models have significantly improved feature representation by integrating information from multiple modalities, making them highly suitable for a broader set of applications. However, the exploration of multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Kaiwen Zheng , Xuri Ge , Junchen Fu , Jun Peng , Joemon M. Jose

Due to the absence of a single standardized imaging protocol, domain shift between data acquired from different sites is an inherent property of medical images and has become a major obstacle for large-scale deployment of learning-based…

Computer Vision and Pattern Recognition · Computer Science 2023-08-15 Dewei Hu , Hao Li , Han Liu , Xing Yao , Jiacheng Wang , Ipek Oguz

With the recent advancement of deep convolutional neural networks, significant progress has been made in general face recognition. However, the state-of-the-art general face recognition models do not generalize well to occluded face images,…

Computer Vision and Pattern Recognition · Computer Science 2022-08-09 Haibo Qiu , Dihong Gong , Zhifeng Li , Wei Liu , Dacheng Tao

Unified multimodal models (UMMs) have emerged as a powerful paradigm in fundamental cross-modality research, demonstrating significant potential in both image understanding and generation. However, existing research in the face domain…

Computer Vision and Pattern Recognition · Computer Science 2026-01-14 Junzhe Li , Sifan Zhou , Liya Guo , Xuerui Qiu , Linrui Xu , Delin Qu , Tingting Long , Chun Fan , Ming Li , Hehe Fan , Jun Liu , Shuicheng Yan