中文
相关论文

相关论文: Zero-Shot Satellite Image Retrieval through Joint …

200 篇论文

Current Synthetic Aperture Radar (SAR)-based flood detection methods face critical limitations that hinder operational deployment. Supervised learning approaches require extensive labeled training data, exhibit poor geographical…

应用统计 · 统计学 2025-10-15 Narumasa Tsutsumida , Tomohiro Tanaka , Nifat Sultana

This paper presents a zero-shot system for fact-checked claim retrieval. We employed several state-of-the-art large language models to obtain text embeddings. The models were then combined to obtain the best possible result. Our approach…

计算与语言 · 计算机科学 2025-08-14 Ladislav Lenc , Daniel Cífka , Jiří Martínek , Jakub Šmíd , Pavel Král

Zero-shot learning (ZSL) aims to recognize objects of novel classes without any training samples of specific classes, which is achieved by exploiting the semantic information and auxiliary datasets. Recently most ZSL approaches focus on…

计算机视觉与模式识别 · 计算机科学 2018-07-25 Huajie Jiang , Ruiping Wang , Shiguang Shan , Xilin Chen

In this paper, we address the problem of global-scale image geolocation, proposing a mixed classification-retrieval scheme. Unlike other methods that strictly tackle the problem as a classification or retrieval task, we combine the two…

计算机视觉与模式识别 · 计算机科学 2021-05-18 Giorgos Kordopatis-Zilos , Panagiotis Galopoulos , Symeon Papadopoulos , Ioannis Kompatsiaris

While the Earth observation community has witnessed a surge in high-impact foundation models and global Earth embedding datasets, a significant barrier remains in translating these academic assets into freely accessible tools. This tutorial…

计算机视觉与模式识别 · 计算机科学 2026-04-09 Yijie Zheng , Weijie Wu , Bingyue Wu , Long Zhao , Guoqing Li , Mikolaj Czerkawski , Konstantin Klemmer

Visual Question Answering (VQA) is a multi-modal task that involves answering questions from an input image, semantically understanding the contents of the image and answering it in natural language. Using VQA for disaster management is an…

计算机视觉与模式识别 · 计算机科学 2022-11-14 Aditya Kane , V Manushree , Sahil Khose

Training a referring expression comprehension (ReC) model for a new visual domain requires collecting referring expressions, and potentially corresponding bounding boxes, for images in the domain. While large-scale pre-trained models are…

计算机视觉与模式识别 · 计算机科学 2022-05-04 Sanjay Subramanian , William Merrill , Trevor Darrell , Matt Gardner , Sameer Singh , Anna Rohrbach

Cross-View Geo-Localization (CVGL) in remote sensing aims to locate a drone-view query by matching it to geo-tagged satellite images. Although supervised methods have achieved strong results on closeset benchmarks, they often fail to…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Jun Lu , Zehao Sang , Haoqi Wei , Xiangyun Liu , Kun Zhu , Haitao Guo , Zhihui Gong , Lei Ding

The Zero-Shot Sketch-based Image Retrieval (ZS-SBIR) is a challenging task because of the large domain gap between sketches and natural images as well as the semantic inconsistency between seen and unseen categories. Previous literature…

计算机视觉与模式识别 · 计算机科学 2022-04-13 Yu-Wei Zhan , Xin Luo , Yongxin Wang , Zhen-Duo Chen , Xin-Shun Xu

Cross-view geo-localization (CVGL) estimates a camera's location by matching a street-view image to geo-referenced overhead imagery, enabling GPS-denied localization and navigation. Existing methods almost universally formulate CVGL as an…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Yunus Talha Erzurumlu , Jiyong Kwag , Alper Yilmaz

The recent growth of large foundation models that can easily generate pseudo-labels for huge quantity of unlabeled data makes unsupervised Zero-Shot Cross-Domain Image Retrieval (UZS-CDIR) less relevant. In this paper, we therefore turn our…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Chor Boon Tan , Conghui Hu , Gim Hee Lee

Stereo foundation models achieve strong zero-shot generalization but remain computationally prohibitive for real-time applications. Efficient stereo architectures, on the other hand, sacrifice robustness for speed and require costly…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Bowen Wen , Shaurya Dewan , Stan Birchfield

Zero-shot composed image retrieval (ZS-CIR) is a rapidly growing area with significant practical applications, allowing users to retrieve a target image by providing a reference image and a relative caption describing the desired…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Yongcong Ye , Kai Zhang , Yanghai Zhang , Enhong Chen , Longfei Li , Jun Zhou

Geographic information is essential for modeling tasks in fields ranging from ecology to epidemiology. However, extracting relevant location characteristics for a given task can be challenging, often requiring expensive data fusion or…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Konstantin Klemmer , Esther Rolf , Caleb Robinson , Lester Mackey , Marc Rußwurm

Human-annotated attributes serve as powerful semantic embeddings in zero-shot learning. However, their annotation process is labor-intensive and needs expert supervision. Current unsupervised semantic embeddings, i.e., word embeddings,…

计算机视觉与模式识别 · 计算机科学 2023-05-29 Wenjia Xu , Yongqin Xian , Jiuniu Wang , Bernt Schiele , Zeynep Akata

Recent advances have demonstrated that Language Vision Models (LVMs) surpass the existing State-of-the-Art (SOTA) in two-dimensional (2D) computer vision tasks, motivating attempts to apply LVMs to three-dimensional (3D) data. While LVMs…

计算机视觉与模式识别 · 计算机科学 2025-09-29 June Moh Goo , Zichao Zeng , Jan Boehm

In the absence of parallax cues, a learning-based single image depth estimation (SIDE) model relies heavily on shading and contextual cues in the image. While this simplicity is attractive, it is necessary to train such models on large and…

计算机视觉与模式识别 · 计算机科学 2024-04-18 Suraj Patni , Aradhye Agarwal , Chetan Arora

Monocular visual SLAM enables 3D reconstruction from internet video and autonomous navigation on resource-constrained platforms, yet suffers from scale drift, i.e., the gradual divergence of estimated scale over long sequences. Existing…

计算机视觉与模式识别 · 计算机科学 2026-01-15 Yuchen Wu , Jiahe Li , Xiaohan Yu , Lina Yu , Jin Zheng , Xiao Bai

Ephemeral gullies are a primary cause of soil erosion and their reliable, accurate, and early detection will facilitate significant improvements in the sustainability of global agricultural systems. In our view, prior research has not…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Seyed Mohamad Ali Tousi , Ramy Farag , Jacket Demby's , Gbenga Omotara , John A. Lory , G. N. DeSouza

Vision-Language Models for remote sensing have shown promising uses thanks to their extensive pretraining. However, their conventional usage in zero-shot scene classification methods still involves dividing large images into patches and…