English
Related papers

Related papers: GEOBIND: Binding Text, Image, and Audio through Sa…

200 papers

Medical vision-language pretraining models (VLPM) have achieved remarkable progress in fusing chest X-rays (CXR) with clinical texts, introducing image-text data binding approaches that enable zero-shot learning and downstream clinical…

Computer Vision and Pattern Recognition · Computer Science 2024-03-21 Yuan Gao , Sangwook Kim , David E Austin , Chris McIntosh

We study the image-based geolocalization problem, aiming to localize ground-view query images on cartographic maps. Current methods often utilize cross-view localization techniques to match ground-view query images with 2D maps. However,…

Computer Vision and Pattern Recognition · Computer Science 2023-11-06 Mengjie Zhou , Liu Liu , Yiran Zhong , Andrew Calway

Remote sensing image fusion is an effective way to use a large volume of data from multisensor images. Most earth satellites such as SPOT, Landsat 7, IKONOS and QuickBird provide both panchromatic (Pan) images at a higher spatial resolution…

Computer Vision and Pattern Recognition · Computer Science 2014-03-24 Reham Gharbia , Ahmad Taher Azar , Ali El Baz , Aboul Ella Hassanien

Despite recent progress in computer vision, fine-grained interpretation of satellite images remains challenging because of a lack of labeled training data. To overcome this limitation, we propose using Wikipedia as a previously untapped…

Computer Vision and Pattern Recognition · Computer Science 2018-09-28 Evan Sheehan , Burak Uzkent , Chenlin Meng , Zhongyi Tang , Marshall Burke , David Lobell , Stefano Ermon

Remote sensing data is crucial for applications ranging from monitoring forest fires and deforestation to tracking urbanization. Most of these tasks require dense pixel-level annotations for the model to parse visual information from…

Computer Vision and Pattern Recognition · Computer Science 2021-10-18 Shasvat Desai , Debasmita Ghose

We introduce a highly multimodal transformer to represent many remote sensing modalities - multispectral optical, synthetic aperture radar, elevation, weather, pseudo-labels, and more - across space and time. These inputs are useful for…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Gabriel Tseng , Anthony Fuller , Marlena Reil , Henry Herzog , Patrick Beukema , Favyen Bastani , James R. Green , Evan Shelhamer , Hannah Kerner , David Rolnick

Estimating the location where an image was taken based solely on the contents of the image is a challenging task, even for humans, as properly labeling an image in such a fashion relies heavily on contextual information, and is not as…

Computer Vision and Pattern Recognition · Computer Science 2017-12-29 Jesse M. Johns , Jeremiah Rounds , Michael J. Henry

In this paper, we address the challenging problem of data association for underwater SLAM through a novel method for sonar image correspondence using learned features. We introduce SONIC (SONar Image Correspondence), a pose-supervised…

Computer Vision and Pattern Recognition · Computer Science 2024-05-15 Samiran Gode , Akshay Hinduja , Michael Kaess

Deep learning has become the gold standard for image processing over the past decade. Simultaneously, we have seen growing interest in orbital activities such as satellite servicing and debris removal that depend on proximity operations…

Machine Learning · Computer Science 2021-01-15 Carson Schubert , Kevin Black , Daniel Fonseka , Abhimanyu Dhir , Jacob Deutsch , Nihal Dhamani , Gavin Martin , Maruthi Akella

Zero-shot learning (ZSL) models rely on learning a joint embedding space where both textual/semantic description of object classes and visual representation of object images can be projected to for nearest neighbour search. Despite the…

Computer Vision and Pattern Recognition · Computer Science 2019-07-22 Li Zhang , Tao Xiang , Shaogang Gong

We capitalize on large amounts of readily-available, synchronous data to learn a deep discriminative representations shared across three major natural modalities: vision, sound and language. By leveraging over a year of sound from video and…

Computer Vision and Pattern Recognition · Computer Science 2017-06-06 Yusuf Aytar , Carl Vondrick , Antonio Torralba

Buildings classification using satellite images is becoming more important for several applications such as damage assessment, resource allocation, and population estimation. We focus, in this work, on buildings damage assessment (BDA) and…

Computer Vision and Pattern Recognition · Computer Science 2021-11-30 Mohammad Dimassi , Abed Ellatif Samhat , Mohammad Zaraket , Jamal Haidar , Mustafa Shukor , Ali J. Ghandour

Heterogeneous collections of ground and airborne imagery can readily be used to create high-quality 3D models and novel viewpoint renderings of the observed scene. Standard photogrammetry pipelines generate models in arbitrary coordinate…

Computer Vision and Pattern Recognition · Computer Science 2025-03-10 Adam Bredvik , Scott Richardson , Daniel Crispell

Matching images and sentences demands a fine understanding of both modalities. In this paper, we propose a new system to discriminatively embed the image and text to a shared visual-textual space. In this field, most existing works apply…

Computer Vision and Pattern Recognition · Computer Science 2021-07-28 Zhedong Zheng , Liang Zheng , Michael Garrett , Yi Yang , Mingliang Xu , Yi-Dong Shen

Deep learning relies heavily on data augmentation to mitigate limited data, especially in medical imaging. Recent multimodal learning integrates text and images for segmentation, known as referring or text-guided image segmentation.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Shurong Chai , Rahul Kumar JAIN , Rui Xu , Shaocong Mo , Ruibo Hou , Shiyu Teng , Jiaqing Liu , Lanfen Lin , Yen-Wei Chen

We propose a novel deep training algorithm for joint representation of audio and visual information which consists of a single stream network (SSNet) coupled with a novel loss function to learn a shared deep latent space representation of…

Computer Vision and Pattern Recognition · Computer Science 2019-09-20 Shah Nawaz , Muhammad Kamran Janjua , Ignazio Gallo , Arif Mahmood , Alessandro Calefati

This paper extends LiDAR-BIND, a modular multi-modal fusion framework that binds heterogeneous sensors (radar, sonar) to a LiDAR-defined latent space, with mechanisms that explicitly enforce temporal consistency. We introduce three…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Niels Balemans , Ali Anwar , Jan Steckel , Siegfried Mercelis

Multi-spectral satellite imaging sensors acquire various spectral band images such as red (R), green (G), blue (B), near-infrared (N), etc. Thanks to the unique spectroscopic property of each spectral band with respective to the objects on…

Image and Video Processing · Electrical Eng. & Systems 2020-02-25 Joonyoung Song , Jae-Heon Jeong , Dae-Soon Park , Hyun-Ho Kim , Doo-Chun Seo , Jong Chul Ye

Interpreting remote sensing imagery enables numerous downstream applications ranging from land-use planning to deforestation monitoring. Robustly classifying this data is challenging due to the Earth's geographic diversity. While many…

Computer Vision and Pattern Recognition · Computer Science 2023-04-25 Jonathan Roberts , Kai Han , Samuel Albanie

The rapid advancement of remote sensing foundation models, particularly vision and multimodal models, has significantly enhanced the capabilities of intelligent geospatial data interpretation. These models combine various data modalities,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Ziyue Huang , Hongxi Yan , Qiqi Zhan , Shuai Yang , Mingming Zhang , Chenkai Zhang , YiMing Lei , Zeming Liu , Qingjie Liu , Yunhong Wang
‹ Prev 1 8 9 10 Next ›