English
Related papers

Related papers: SatBLIP: Context Understanding and Feature Identif…

200 papers

We study the task of locating a user in a mapped indoor environment using natural language queries and images from the environment. Building on recent pretrained vision-language models, we learn a similarity score between text descriptions…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Seth Pate , Lawson L. S. Wong

Automated crop mapping through Satellite Image Time Series (SITS) has emerged as a crucial avenue for agricultural monitoring and management. However, due to the low resolution and unclear parcel boundaries, annotating pixel-level masks is…

Computer Vision and Pattern Recognition · Computer Science 2024-12-06 Hao Zhu , Yan Zhu , Jiayu Xiao , Tianxiang Xiao , Yike Ma , Yucheng Zhang , Feng Dai

Satellite-based slum segmentation holds significant promise in generating global estimates of urban poverty. However, the morphological heterogeneity of informal settlements presents a major challenge, hindering the ability of models…

Computer Vision and Pattern Recognition · Computer Science 2025-11-14 Sumin Lee , Sungwon Park , Jeasurk Yang , Jihee Kim , Meeyoung Cha

Understanding the semantics of individual regions or patches of unconstrained images, such as open-world object detection, remains a critical yet challenging task in computer vision. Building on the success of powerful image-level…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Haosen Yang , Chuofan Ma , Bin Wen , Yi Jiang , Zehuan Yuan , Xiatian Zhu

The objective in this paper is to improve the performance of text-to-image retrieval. To this end, we introduce a new framework that can boost the performance of large-scale pre-trained vision-language models, so that they can be used for…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Guanqi Zhan , Yuanpei Liu , Kai Han , Weidi Xie , Andrew Zisserman

Urban region profiling from web-sourced data is of utmost importance for urban planning and sustainable development. We are witnessing a rising trend of LLMs for various fields, especially dealing with multi-modal data research such as…

Computation and Language · Computer Science 2024-03-26 Yibo Yan , Haomin Wen , Siru Zhong , Wei Chen , Haodong Chen , Qingsong Wen , Roger Zimmermann , Yuxuan Liang

Situational Graphs (S-Graphs) merge geometric models of the environment generated by Simultaneous Localization and Mapping (SLAM) approaches with 3D scene graphs into a multi-layered jointly optimizable factor graph. As an advantage,…

CLIP is a widely used foundational vision-language model that is used for zero-shot image recognition and other image-text alignment tasks. We demonstrate that CLIP is vulnerable to change in image quality under compression. This surprising…

Computer Vision and Pattern Recognition · Computer Science 2023-11-27 Cangxiong Chen , Vinay P. Namboodiri , Julian Padget

Change detection, which typically relies on the comparison of bi-temporal images, is significantly hindered when only a single image is available. Comparing a single image with an existing map, such as OpenStreetMap, which is continuously…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Shuguo Jiang , Fang Xu , Sen Jia , Gui-Song Xia

Image captioning, a fundamental task in vision-language understanding, seeks to generate accurate natural language descriptions for provided images. Current image captioning approaches heavily rely on high-quality image-caption pairs, which…

Computer Vision and Pattern Recognition · Computer Science 2023-11-03 Chuanyang Jin

Camera traps offer enormous new opportunities in ecological studies, but current automated image analysis methods often lack the contextual richness needed to support impactful conservation outcomes. Here we present an integrated approach…

Computer Vision and Pattern Recognition · Computer Science 2024-11-22 Paul Fergus , Carl Chalmers , Naomi Matthews , Stuart Nixon , Andre Burger , Oliver Hartley , Chris Sutherland , Xavier Lambin , Steven Longmore , Serge Wich

In this paper we address the challenge of land cover classification for satellite images via Deep Learning (DL). Land Cover aims to detect the physical characteristics of the territory and estimate the percentage of land occupied by a…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Eleonora Bernasconi , Francesco Pugliese , Diego Zardetto , Monica Scannapieco

While significant progress has been made towards explaining black-box machine-learning (ML) models, there is still a distinct lack of diagnostic tools that elucidate the spatial behaviour of ML models in terms of predictive skill and…

Machine Learning · Computer Science 2023-06-01 Alexander Brenning

Guide dog robots offer promising solutions to enhance mobility and safety for visually impaired individuals, addressing the limitations of traditional guide dogs, particularly in perceptual intelligence and communication. With the emergence…

Robotics · Computer Science 2025-02-13 ByungOk Han , Woo-han Yun , Beom-Su Seo , Jaehong Kim

Deployment of robots into hazardous environments typically involves a ``Human-Robot Teaming'' (HRT) paradigm, in which a human supervisor interacts with a remotely operating robot inside the hazardous zone. Situational Awareness (SA) is…

In this study, we define and tackle zero shot "real" classification by description, a novel task that evaluates the ability of Vision-Language Models (VLMs) like CLIP to classify objects based solely on descriptive attributes, excluding…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Ethan Baron , Idan Tankel , Peter Tu , Guy Ben-Yosef

Humans interpret safety not as a binary signal but as a continuous, context- and spatially-dependent notion of risk. While risk is subjective, humans form rational mental models that guide action selection in dynamic environments. This work…

Robotics · Computer Science 2025-12-10 Timothy Chen , Marcus Dominguez-Kuhne , Aiden Swann , Xu Liu , Mac Schwager

The ability to develop a high-level understanding of a scene, such as perceiving danger levels, can prove valuable in planning multi-robot search and rescue (SaR) missions. In this work, we propose to uniquely leverage natural language…

Robotics · Computer Science 2021-04-09 Vikram Shree , Beatriz Asfora , Rachel Zheng , Samantha Hong , Jacopo Banfi , Mark Campbell

Remote sensing image plays an irreplaceable role in fields such as agriculture, water resources, military, and disaster relief. Pixel-level interpretation is a critical aspect of remote sensing image applications; however, a prevalent…

Computer Vision and Pattern Recognition · Computer Science 2024-11-05 Kaiyu Li , Ruixun Liu , Xiangyong Cao , Xueru Bai , Feng Zhou , Deyu Meng , Zhi Wang

Dense visual perception tasks have been constrained by their reliance on predefined categories, limiting their applicability in real-world scenarios where visual concepts are unbounded. While Vision-Language Models (VLMs) like CLIP have…

Computer Vision and Pattern Recognition · Computer Science 2025-08-18 Junjie Wang , Keyu Chen , Yulin Li , Bin Chen , Hengshuang Zhao , Xiaojuan Qi , Zhuotao Tian
‹ Prev 1 3 4 5 6 7 10 Next ›