English
Related papers

Related papers: Hierarchical Cross-Modal Alignment for Open-Vocabu…

200 papers

In autonomous driving, 3D object detection provides more precise information for downstream tasks, including path planning and motion estimation, compared to 2D object detection. In this paper, we propose SeSame: a method aimed at enhancing…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Hayeon O , Chanuk Yang , Kunsoo Huh

Despite advances in object detection, aerial imagery remains a challenging domain, as models often fail to generalize across variations in spatial resolution, scene composition, and semantic label coverage. Differences in geographic…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Pourya Shamsolmoali , Masoumeh Zareapoor , Michael Felsberg , Nick Pears , Yue Lu

In this paper, we present a hierarchical question-answering (QA) approach for scene understanding in autonomous vehicles, balancing cost-efficiency with detailed visual interpretation. The method fine-tunes a compact vision-language model…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Safaa Abdullahi Moallim Mohamud , Minjin Baek , Dong Seog Han

The visual commonsense reasoning (VCR) task is to choose an answer and provide a justifying rationale based on the given image and textural question. Representative works first recognize objects in images and then associate them with key…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Jian Zhu , Hanli Wang , Miaojing Shi

Human-object interaction (HOI) detection has seen advancements with Vision Language Models (VLMs), but these methods often depend on extensive manual annotations. Vision Large Language Models (VLLMs) can inherently recognize and reason…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Jianjun Gao , Chen Cai , Ruoyu Wang , Wenyang Liu , Kim-Hui Yap , Kratika Garg , Boon-Siew Han

Open-world instance-level scene understanding aims to locate and recognize unseen object categories that are not present in the annotated dataset. This task is challenging because the model needs to both localize novel 3D objects and infer…

Computer Vision and Pattern Recognition · Computer Science 2023-08-02 Runyu Ding , Jihan Yang , Chuhui Xue , Wenqing Zhang , Song Bai , Xiaojuan Qi

Large foundation models trained on large-scale vision-language data can boost Open-Vocabulary Object Detection (OVD) via synthetic training data, yet the hand-crafted pipelines often introduce bias and overfit to specific prompts. We…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Yang Zhou , Shiyu Zhao , Yuxiao Chen , Zhenting Wang , Can Jin , Dimitris N. Metaxas

Utilizing vision and language models (VLMs) pre-trained on large-scale image-text pairs is becoming a promising paradigm for open-vocabulary visual recognition. In this work, we extend this paradigm by leveraging motion and audio that…

Computer Vision and Pattern Recognition · Computer Science 2022-07-18 Rui Qian , Yeqing Li , Zheng Xu , Ming-Hsuan Yang , Serge Belongie , Yin Cui

Recently, Multimodal Large Language Models (MLLMs) have demonstrated impressive performance on instruction-following tasks by integrating pretrained visual encoders with large language models (LLMs). However, existing approaches often…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Wayner Barrios , Andrés Villa , Juan León Alcázar , SouYoung Jin , Bernard Ghanem

We introduce InteractVLM, a novel method to estimate 3D contact points on human bodies and objects from single in-the-wild images, enabling accurate human-object joint reconstruction in 3D. This is challenging due to occlusions, depth…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Sai Kumar Dwivedi , Dimitrije Antić , Shashank Tripathi , Omid Taheri , Cordelia Schmid , Michael J. Black , Dimitrios Tzionas

Despite the remarkable progress in open-vocabulary object detection (OVD), a significant gap remains between the training and testing phases. During training, the RPN and RoI heads often misclassify unlabeled novel-category objects as…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Yupeng Zhang , Ruize Han , Zhiwei Chen , Wei Feng , Liang Wan

Multimodal fusion of remote sensing images serves as a core technology for overcoming the limitations of single-source data and improving the accuracy of surface information extraction, which exhibits significant application value in fields…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Siyu Zhang , Lianlei Shan , Runhe Qiu

This work presents OVIR-3D, a straightforward yet effective method for open-vocabulary 3D object instance retrieval without using any 3D data for training. Given a language query, the proposed method is able to return a ranked set of 3D…

Computer Vision and Pattern Recognition · Computer Science 2023-11-07 Shiyang Lu , Haonan Chang , Eric Pu Jing , Abdeslam Boularias , Kostas Bekris

Recent advancements in autonomous driving, augmented reality, robotics, and embodied intelligence have necessitated 3D perception algorithms. However, current 3D perception methods, especially specialized small models, exhibit poor…

Computer Vision and Pattern Recognition · Computer Science 2025-02-14 Fan Yang , Sicheng Zhao , Yanhao Zhang , Hui Chen , Haonan Lu , Jungong Han , Guiguang Ding

Pre-trained vision-language models (VLMs) learn to align vision and language representations on large-scale datasets, where each image-text pair usually contains a bag of semantic concepts. However, existing open-vocabulary object detectors…

Computer Vision and Pattern Recognition · Computer Science 2023-02-28 Size Wu , Wenwei Zhang , Sheng Jin , Wentao Liu , Chen Change Loy

Existing alignment techniques for Large Language Models (LLMs), such as Direct Preference Optimization (DPO), typically treat the model as a monolithic entity, applying uniform optimization pressure across all layers. This approach…

Computation and Language · Computer Science 2025-10-15 Yukun Zhang , Qi Dong

There is a gap in the understanding of occluded objects in existing large-scale visual language multi-modal models. Current state-of-the-art multi-modal models fail to provide satisfactory results in describing occluded objects through…

Computer Vision and Pattern Recognition · Computer Science 2024-10-04 Shuxin Yang , Xinhan Di

Rare-object detection remains a challenging task in autonomous driving systems, particularly when relying solely on point cloud data. Although Vision-Language Models (VLMs) exhibit strong capabilities in image understanding, their potential…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Mai Tsujimoto

Open-vocabulary detection (OVD) is an object detection task aiming at detecting objects from novel categories beyond the base categories on which the detector is trained. Recent OVD methods rely on large-scale visual-language pre-trained…

Computer Vision and Pattern Recognition · Computer Science 2023-03-24 Xiaoshi Wu , Feng Zhu , Rui Zhao , Hongsheng Li

Single-modal object detection tasks often experience performance degradation when encountering diverse scenarios. In contrast, multimodal object detection tasks can offer more comprehensive information about object features by integrating…

Computer Vision and Pattern Recognition · Computer Science 2024-12-25 Chang Liu , Xin Ma , Xiaochen Yang , Yuxiang Zhang , Yanni Dong