English
Related papers

Related papers: When Language Model Guides Vision: Grounding DINO …

200 papers

The Mixture of Experts (MoE) architecture has excelled in Large Vision-Language Models (LVLMs), yet its potential in real-time open-vocabulary object detectors, which also leverage large-scale vision-language datasets but smaller models,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-24 Yehao Lu , Minghe Weng , Zekang Xiao , Rui Jiang , Wei Su , Guangcong Zheng , Ping Lu , Xi Li

We introduce a multimodal vision framework for precision livestock farming, harnessing the power of GroundingDINO, HQSAM, and ViTPose models. This integrated suite enables comprehensive behavioral analytics from video data without invasive…

Computer Vision and Pattern Recognition · Computer Science 2024-06-17 Ahmed Qazi , Taha Razzaq , Asim Iqbal

Deep learning-based object detection has revolutionized Precision Livestock Farming (PLF), yet a critical barrier remains: high-performance Foundation Models (such as SAM 3) are too computationally intensive for edge deployment, while…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Marcos Vinicius Mendes Faria , Thiago Borges Pereira , Isabella C. F. S. Condotta , Thiago Meireles Paixão , Francisco de Assis Boldt

Cattle identification is critical for efficient livestock farming management, currently reliant on radio-frequency identification (RFID) ear tags. However, RFID-based systems are prone to failure due to loss, damage, tampering, and…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Rabin Dulal , Lihong Zheng , Ashad Kabir

Monitoring feeding behaviour is a relevant task for efficient herd management and the effective use of available resources in grazing cattle. The ability to automatically recognise animals' feeding activities through the identification of…

In this paper, we introduce DINO-X, which is a unified object-centric vision model developed by IDEA Research with the best open-world object detection performance to date. DINO-X employs the same Transformer-based encoder-decoder…

In the poultry industry, detecting chicken illnesses is essential to avoid financial losses. Conventional techniques depend on manual observation, which is laborious and prone to mistakes. Using YOLO v8 a deep learning model for real-time…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Akhil Saketh Reddy Sabbella , Ch. Lakshmi Prachothan , Eswar Kumar Panta

Recent breakthroughs in large foundation models have enabled the possibility of transferring knowledge pre-trained on vast datasets to domains with limited data availability. Agriculture is one of the domains that lacks sufficient data.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Yanan Wang , Zhenghao Fei , Ruichen Li , Yibin Ying

Open-vocabulary object detection with vision-language models (VLMs) such as Grounding DINO suffers from performance degradation under test-time distribution shifts, primarily due to semantic misalignment between text embeddings and shifted…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Lihua Zhou , Mao Ye , Xiatian Zhu , Nianxin Li , Changyi Ma , Shuaifeng Li , Yitong Qin , Hongbin Liu , Jiebo Luo , Zhen Lei

Recent approaches have shown that training deep neural networks directly on large-scale image-text pair collections enables zero-shot transfer on various recognition tasks. One central issue is how this can be generalized to object…

Computer Vision and Pattern Recognition · Computer Science 2022-08-30 Johnathan Xie , Shuai Zheng

Accurate and generalizable object segmentation in ultrasound imaging remains a significant challenge due to anatomical variability, diverse imaging protocols, and limited annotated data. In this study, we propose a prompt-driven…

Computer Vision and Pattern Recognition · Computer Science 2025-09-10 Hamza Rasaee , Taha Koleilat , Hassan Rivaz

We present Neural Congealing -- a zero-shot self-supervised framework for detecting and jointly aligning semantically-common content across a given set of images. Our approach harnesses the power of pre-trained DINO-ViT features to learn:…

Computer Vision and Pattern Recognition · Computer Science 2023-03-07 Dolev Ofri-Amar , Michal Geyer , Yoni Kasten , Tali Dekel

Multi-animal tracking is crucial for understanding animal ecology and behavior. However, it remains a challenging task due to variations in habitat, motion patterns, and species appearance. Traditional approaches typically require extensive…

Computer Vision and Pattern Recognition · Computer Science 2025-11-05 Jan Frederik Meier , Timo Lüddecke

Current large multimodal models (LMMs) face challenges in grounding, which requires the model to relate language components to visual entities. Contrary to the common practice that fine-tunes LMMs with additional grounding supervision, we…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Shengcao Cao , Liang-Yan Gui , Yu-Xiong Wang

Deploying autonomous edge robotics in dynamic military environments is constrained by both scarce domain-specific training data and the computational limits of edge hardware. This paper introduces a hierarchical, zero-shot framework that…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Jesse Barkley , Abraham George , Amir Barati Farimani

Open-vocabulary object detectors such as Grounding DINO are trained on vast and diverse data, achieving remarkable performance on challenging datasets. Due to that, it is unclear where to find their limitations, which is of major concern…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Annika Mütze , Sadia Ilyas , Christian Dörpelkus , Matthias Rottmann

Masked language models like BERT can perform text classification in a zero-shot fashion by reformulating downstream tasks as text infilling. However, this approach is highly sensitive to the template used to prompt the model, yet…

Computation and Language · Computer Science 2022-10-27 Mozes van de Kar , Mengzhou Xia , Danqi Chen , Mikel Artetxe

Large-scale vision-language models (VLMs), such as CLIP, have achieved remarkable success in zero-shot learning (ZSL) by leveraging large-scale visual-text pair datasets. However, these methods often lack interpretability, as they compute…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Shiming Chen , Bowen Duan , Salman Khan , Fahad Shahbaz Khan

In recent years, deep learning models have become the standard for agricultural computer vision. Such models are typically fine-tuned to agricultural tasks using model weights that were originally fit to more general, non-agricultural…

Computer Vision and Pattern Recognition · Computer Science 2022-08-05 Amogh Joshi , Dario Guevara , Mason Earles

Bird image segmentation remains a challenging task in computer vision due to extreme pose diversity, complex plumage patterns, and variable lighting conditions. This paper presents a dual-pipeline framework for binary bird image…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Abhinav Munagala