English
Related papers

Related papers: DINO Eats CLIP: Adapting Beyond Knowns for Open-se…

200 papers

Continual learning of vision-language models (VLMs) focuses on leveraging cross-modal pretrained knowledge to incrementally adapt to expanding downstream tasks and datasets, while tackling the challenge of knowledge forgetting. Existing…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Chiyuan He , Zihuan Qiu , Fanman Meng , Linfeng Xu , Qingbo Wu , Hongliang Li

Multi-label classification is crucial for comprehensive image understanding, yet acquiring accurate annotations is challenging and costly. To address this, a recent study suggests exploiting unsupervised multi-label classification…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Dongseob Kim , Hyunjung Shim

Recent work has explored how individual components of the CLIP-ViT model contribute to the final representation by leveraging the shared image-text representation space of CLIP. These components, such as attention heads and MLPs, have been…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Sriram Balasubramanian , Samyadeep Basu , Soheil Feizi

Diffusion models have shown exceptional performance in visual generation tasks. Recently, these models have shifted from traditional U-Shaped CNN-Attention hybrid structures to fully transformer-based isotropic architectures. While these…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Yuchuan Tian , Jing Han , Chengcheng Wang , Yuchen Liang , Chao Xu , Hanting Chen

Face recognition systems are increasingly used in biometric security for convenience and effectiveness. However, they remain vulnerable to spoofing attacks, where attackers use photos, videos, or masks to impersonate legitimate users. This…

Computer Vision and Pattern Recognition · Computer Science 2024-10-23 Arman Keresh , Pakizar Shamoi

Numerous methods have been proposed to adapt a pre-trained foundational CLIP model for few-shot classification. As CLIP is trained on a large corpus, it generalises well through adaptation to few-shot classification. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2024-09-18 Alexey Kravets , Vinay Namboodiri

Although CLIP-like Visual Language Models provide a functional joint feature space for image and text, due to the limitation of the CILP-like model's image input size (e.g., 224), subtle details are lost in the feature representation if we…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Zilun Zhang , Cuifeng Shen , Yuan Shen , Xinyu Zhou , Huixin Xiong , Tiancheng Zhao , Jianwei Yin

Open-vocabulary object detection (OVOD) aims to recognize novel objects whose categories are not included in the training set. In order to classify these unseen classes during training, many OVOD frameworks leverage the zero-shot capability…

Computer Vision and Pattern Recognition · Computer Science 2024-02-22 Joonhyun Jeong , Geondo Park , Jayeon Yoo , Hyungsik Jung , Heesu Kim

Current self-supervised monocular depth estimation (MDE) approaches encounter performance limitations due to insufficient semantic-spatial knowledge extraction. To address this challenge, we propose Hybrid-depth, a novel framework that…

Computer Vision and Pattern Recognition · Computer Science 2025-10-13 Wenyao Zhang , Hongsi Liu , Bohan Li , Jiawei He , Zekun Qi , Yunnan Wang , Shengyang Zhao , Xinqiang Yu , Wenjun Zeng , Xin Jin

While Contrastive Language-Image Pre-training (CLIP) has advanced open-vocabulary predictions, its performance on semantic segmentation remains suboptimal. This shortfall primarily stems from its spatial-invariant semantic features and…

Computer Vision and Pattern Recognition · Computer Science 2024-11-15 Yuheng Shi , Minjing Dong , Chang Xu

Unsupervised domain adaptation (UDA) has proven to be very effective in transferring knowledge obtained from a source domain with labeled data to a target domain with unlabeled data. Owing to the lack of labeled data in the target domain…

Computer Vision and Pattern Recognition · Computer Science 2023-08-01 Qing Yu , Go Irie , Kiyoharu Aizawa

In this paper, we present an open-set object detector, called Grounding DINO, by marrying Transformer-based detector DINO with grounded pre-training, which can detect arbitrary objects with human inputs such as category names or referring…

Computer Vision and Pattern Recognition · Computer Science 2024-07-22 Shilong Liu , Zhaoyang Zeng , Tianhe Ren , Feng Li , Hao Zhang , Jie Yang , Qing Jiang , Chunyuan Li , Jianwei Yang , Hang Su , Jun Zhu , Lei Zhang

Contrastive Language-Image Pre-training (CLIP) has made a remarkable breakthrough in open-vocabulary zero-shot image recognition. Many recent studies leverage the pre-trained CLIP models for image-level classification and manipulation. In…

Computer Vision and Pattern Recognition · Computer Science 2022-07-28 Chong Zhou , Chen Change Loy , Bo Dai

The performance of CLIP in dynamic facial expression recognition (DFER) task doesn't yield exceptional results as observed in other CLIP-based classification tasks. While CLIP's primary objective is to achieve alignment between images and…

Computer Vision and Pattern Recognition · Computer Science 2024-03-08 Zeng Tao , Yan Wang , Junxiong Lin , Haoran Wang , Xinji Mai , Jiawen Yu , Xuan Tong , Ziheng Zhou , Shaoqi Yan , Qing Zhao , Liyuan Han , Wenqiang Zhang

Accurate 3D scene reconstruction is essential for numerous medical tasks. Given the challenges in obtaining ground truth data, there has been an increasing focus on self-supervised learning (SSL) for endoscopic depth estimation as a basis…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Beilei Cui , Long Bai , Mobarakol Islam , An Wang , Zhiqi Ma , Yiming Huang , Feng Li , Zhen Chen , Zhongliang Jiang , Nassir Navab , Hongliang Ren

Visual-language models such as CLIP provide powerful general-purpose representations, but their raw embeddings are not optimized for supervised classification, often exhibiting limited class separation and excessive dimensionality. We…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Matej Suchanek , Klara Janouskova , Ondrej Vasatko , Jiri Matas

In this paper, we focus on unsupervised learning for Video Object Segmentation (VOS) which learns visual correspondence (i.e., the similarity between pixel-level features) from unlabeled videos. Previous methods are mainly based on the…

Computer Vision and Pattern Recognition · Computer Science 2022-10-25 Xiao Pan , Peike Li , Zongxin Yang , Huiling Zhou , Chang Zhou , Hongxia Yang , Jingren Zhou , Yi Yang

Robotic and autonomous systems need dense spatial cues, but many monocular depth models are heavy, task-specific, or hard to attach to an existing multimodal stack. CLIP offers strong semantic representations, yet most CLIP-based depth…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Taewan Cho , Taeryang Kim , Andrew Jaeyong Choi

In indoor scenes, the diverse distribution of object locations and scales makes the visual 3D perception task a big challenge. Previous works (e.g, NeRF-Det) have demonstrated that implicit representation has the capacity to benefit the…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Chi Huang , Xinyang Li , Yansong Qu , Changli Wu , Xiaofan Li , Shengchuan Zhang , Liujuan Cao

View based strategies for 3D object recognition have proven to be very successful. The state-of-the-art methods now achieve over 90% correct category level recognition performance on appearance images. We improve upon these methods by…

Computer Vision and Pattern Recognition · Computer Science 2019-06-05 Chu Wang , Marcello Pelillo , Kaleem Siddiqi