English
Related papers

Related papers: Grounding DINO: Marrying DINO with Grounded Pre-Tr…

200 papers

Open-vocabulary object detection with vision-language models (VLMs) such as Grounding DINO suffers from performance degradation under test-time distribution shifts, primarily due to semantic misalignment between text embeddings and shifted…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Lihua Zhou , Mao Ye , Xiatian Zhu , Nianxin Li , Changyi Ma , Shuaifeng Li , Yitong Qin , Hongbin Liu , Jiebo Luo , Zhen Lei

This paper presents DINO-RotateMatch, a deep-learning framework designed to address the chal lenges of image matching in large-scale 3D reconstruction from unstructured Internet images. The method integrates a dataset-adaptive image pairing…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Kaichen Zhang , Tianxiang Sheng , Xuanming Shi

Localizing objects in 3D scenes according to the semantics of a given natural language is a fundamental yet important task in the field of multimedia understanding, which benefits various real-world applications such as robotics and…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Wencan Huang , Daizong Liu , Wei Hu

Recent approaches have shown that training deep neural networks directly on large-scale image-text pair collections enables zero-shot transfer on various recognition tasks. One central issue is how this can be generalized to object…

Computer Vision and Pattern Recognition · Computer Science 2022-08-30 Johnathan Xie , Shuai Zheng

Object detection in civil engineering applications is constrained by limited annotated data in specialized domains. We introduce DINO-YOLO, a hybrid architecture combining YOLOv12 with DINOv3 self-supervised vision transformers for…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Malaisree P , Youwai S , Kitkobsin T , Janrungautai S , Amorndechaphon D , Rojanavasu P

Phrase grounding models localize an object in the image given a referring expression. The annotated language queries available during training are limited, which also limits the variations of language combinations that a model can see…

Computer Vision and Pattern Recognition · Computer Science 2020-11-06 Haidong Zhu , Arka Sadhu , Zhaoheng Zheng , Ram Nevatia

This paper presents INGRESS, a robot system that follows human natural language instructions to pick and place everyday objects. The core issue here is the grounding of referring expressions: infer objects and their relationships from input…

Robotics · Computer Science 2018-06-12 Mohit Shridhar , David Hsu

The connection between our 3D surroundings and the descriptive language that characterizes them would be well-suited for localizing and generating human motion in context but for one problem. The complexity introduced by multiple modalities…

Computer Vision and Pattern Recognition · Computer Science 2024-05-30 Zoltán Á. Milacski , Koichiro Niinuma , Ryosuke Kawamura , Fernando de la Torre , László A. Jeni

When trained at a sufficient scale, self-supervised learning has exhibited a notable ability to solve a wide range of visual or language understanding tasks. In this paper, we investigate simple, yet effective approaches for adapting the…

Computer Vision and Pattern Recognition · Computer Science 2022-10-28 Chaofan Ma , Yuhuan Yang , Yanfeng Wang , Ya Zhang , Weidi Xie

Visual grounding, a crucial vision-language task involving the understanding of the visual context based on the query expression, necessitates the model to capture the interactions between objects, as well as various spatial and attribute…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Haozhan Shen , Tiancheng Zhao , Mingwei Zhu , Jianwei Yin

The goal of this paper is to improve the generality and accuracy of open-vocabulary object counting in images. To improve the generality, we repurpose an open-vocabulary detection foundation model (GroundingDINO) for the counting task, and…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Niki Amini-Naieni , Tengda Han , Andrew Zisserman

Object-centric understanding is fundamental to human vision and required for complex reasoning. Traditional methods define slot-based bottlenecks to learn object properties explicitly, while recent self-supervised vision models like DINO…

Computer Vision and Pattern Recognition · Computer Science 2025-10-03 Stefan Sylvius Wagner , Stefan Harmeling

The current state-of-the-art methods in domain adaptive object detection (DAOD) use Mean Teacher self-labelling, where a teacher model, directly derived as an exponential moving average of the student model, is used to generate labels on…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Marc-Antoine Lavoie , Anas Mahmoud , Steven L. Waslander

Existing object detection methods are bounded in a fixed-set vocabulary by costly labeled data. When dealing with novel categories, the model has to be retrained with more bounding box annotations. Natural language supervision is an…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Chuang Lin , Peize Sun , Yi Jiang , Ping Luo , Lizhen Qu , Gholamreza Haffari , Zehuan Yuan , Jianfei Cai

Accurate and generalizable object segmentation in ultrasound imaging remains a significant challenge due to anatomical variability, diverse imaging protocols, and limited annotated data. In this study, we propose a prompt-driven…

Computer Vision and Pattern Recognition · Computer Science 2025-09-10 Hamza Rasaee , Taha Koleilat , Hassan Rivaz

Open-set object detection (OSOD) is highly desirable for robotic manipulation in unstructured environments. However, existing OSOD methods often fail to meet the requirements of robotic applications due to their high computational burden…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Yonghao He , Hu Su , Haiyong Yu , Cong Yang , Wei Sui , Cong Wang , Song Liu

Many artwork collections contain textual attributes that provide rich and contextualised descriptions of artworks. Visual grounding offers the potential for localising subjects within these descriptions on images, however, existing…

Computer Vision and Pattern Recognition · Computer Science 2024-10-17 Selina Khan , Nanne van Noord

Intention-oriented object detection aims to detect desired objects based on specific intentions or requirements. For instance, when we desire to "lie down and rest", we instinctively seek out a suitable option such as a "bed" or a "sofa"…

Computer Vision and Pattern Recognition · Computer Science 2023-10-27 Mengxue Qu , Yu Wu , Wu Liu , Xiaodan Liang , Jingkuan Song , Yao Zhao , Yunchao Wei

We propose FindIt, a simple and versatile framework that unifies a variety of visual grounding and localization tasks including referring expression comprehension, text-based localization, and object detection. Key to our architecture is an…

Computer Vision and Pattern Recognition · Computer Science 2022-08-10 Weicheng Kuo , Fred Bertsch , Wei Li , AJ Piergiovanni , Mohammad Saffar , Anelia Angelova

Identifying defects and anomalies in industrial products is a critical quality control task. Traditional manual inspection methods are slow, subjective, and error-prone. In this work, we propose a novel zero-shot training-free approach for…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Tsun-Hin Cheung , Ka-Chun Fung , Songjiang Lai , Kwan-Ho Lin , Vincent Ng , Kin-Man Lam