English
Related papers

Related papers: GEOBIND: Binding Text, Image, and Audio through Sa…

200 papers

We present TaxaBind, a unified embedding space for characterizing any species of interest. TaxaBind is a multimodal embedding space across six modalities: ground-level images of species, geographic location, satellite image, text, audio,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-04 Srikumar Sastry , Subash Khanal , Aayush Dhakal , Adeel Ahmad , Nathan Jacobs

We present ImageBind, an approach to learn a joint embedding across six different modalities - images, text, audio, depth, thermal, and IMU data. We show that all combinations of paired data are not necessary to train such a joint…

Computer Vision and Pattern Recognition · Computer Science 2023-06-01 Rohit Girdhar , Alaaeldin El-Nouby , Zhuang Liu , Mannat Singh , Kalyan Vasudev Alwala , Armand Joulin , Ishan Misra

Cross-view geo-localization infers a location by retrieving geo-tagged reference images that visually correspond to a query image. However, the traditional satellite-centric paradigm limits robustness when high-resolution or up-to-date…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Zixuan Song , Jing Zhang , Di Wang , Zidie Zhou , Wenbin Liu , Haonan Guo , En Wang , Bo Du

Recent image-to-audio models have shown impressive performance on object-centric visual scenes. However, their application to satellite imagery remains limited by the complex, wide-area semantic ambiguity of top-down views. While satellite…

Multimedia · Computer Science 2026-04-17 Kunlin Wu , Yanning Wang , Haofeng Tan , Boyi Chen , Teng Fei , Xianping Ma , Yang Yue , Zan Zhou , Xiaofeng Liu

We simplify space binding by focusing on two core components, a single encoder per modality and high-quality data; enabling training state-of-the-art models on a single GPU in a few hours as opposed to multiple days. We present EBind, an…

Machine Learning · Computer Science 2025-11-19 Jim Broadbent , Felix Cohen , Frederik Hvilshøj , Eric Landau , Eren Sasoglu

Unified multi-model representation spaces are the foundation of multimodal understanding and generation. However, the billions of model parameters and catastrophic forgetting problems make it challenging to further enhance pre-trained…

Computer Vision and Pattern Recognition · Computer Science 2024-05-13 Zehan Wang , Ziang Zhang , Xize Cheng , Rongjie Huang , Luping Liu , Zhenhui Ye , Haifeng Huang , Yang Zhao , Tao Jin , Peng Gao , Zhou Zhao

We propose a vision-based method that localizes a ground vehicle using publicly available satellite imagery as the only prior knowledge of the environment. Our approach takes as input a sequence of ground-level images acquired by the…

Robotics · Computer Science 2022-03-08 Dong-Ki Kim , Matthew R. Walter

We present Sat2Sound, a unified multimodal framework for geospatial soundscape understanding, designed to predict and map the distribution of sounds across the Earth's surface. Existing methods for this task rely on paired satellite images…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Subash Khanal , Srikumar Sastry , Aayush Dhakal , Adeel Ahmad , Abby Stylianou , Nathan Jacobs

We present GeoSynth, a model for synthesizing satellite images with global style and image-driven layout control. The global style control is via textual prompts or geographic location. These enable the specification of scene semantics or…

Computer Vision and Pattern Recognition · Computer Science 2024-04-11 Srikumar Sastry , Subash Khanal , Aayush Dhakal , Nathan Jacobs

A large variety of geospatial data layers is available around the world ranging from remotely-sensed raster data like satellite imagery, digital elevation models, predicted land cover maps, and human-annotated data, to data derived from…

Computer Vision and Pattern Recognition · Computer Science 2025-07-21 Arjun Rao , Esther Rolf

Combining satellite imagery with machine learning (SIML) has the potential to address global challenges by remotely estimating socioeconomic and environmental conditions in data-poor regions, yet the resource requirements of SIML limit its…

We present UniBind, a flexible and efficient approach that learns a unified representation space for seven diverse modalities -- images, text, audio, point cloud, thermal, video, and event data. Existing works, eg., ImageBind, treat the…

Computer Vision and Pattern Recognition · Computer Science 2024-03-20 Yuanhuiyi Lyu , Xu Zheng , Jiazhou Zhou , Lin Wang

Environmental soundscapes convey substantial ecological and social information regarding urban environments; however, their potential remains largely untapped in large-scale geographic analysis. In this study, we investigate the extent to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Pengyu Chen , Xiao Huang , Teng Fei , Sicheng Wang

Recently, human-computer interaction with various modalities has shown promising applications, like GPT-4o and Gemini. Given the foundational role of multimodal joint representation in understanding and generation pipelines, high-quality…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Zehan Wang , Ziang Zhang , Hang Zhang , Luping Liu , Rongjie Huang , Xize Cheng , Hengshuang Zhao , Zhou Zhao

Publicly available satellite imagery, such as Sentinel- 2, often lacks the spatial resolution required for accurate analysis of remote sensing tasks including urban planning and disaster response. Current super-resolution techniques are…

Computer Vision and Pattern Recognition · Computer Science 2025-01-31 Daniel Panangian , Ksenia Bittner

Visual transformers have driven major progress in remote sensing image analysis, particularly in object detection and segmentation. Recent vision-language and multimodal models further extend these capabilities by incorporating auxiliary…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Yu Li , Guilherme N. DeSouza , Praveen Rao , Chi-Ren Shyu

Recent advances in deep-learning based methods for image matching have demonstrated their superiority over traditional algorithms, enabling correspondence estimation in challenging scenes with significant differences in viewing angles,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Rahul Deshmukh , Avinash Kak

In recent years, with the development of aerospace technology, we use more and more images captured by satellites to obtain information. But a large number of useless raw images, limited data storage resource and poor transmission…

Computer Vision and Pattern Recognition · Computer Science 2019-12-11 Junxing Hu , Ling Li , Yijun Lin , Fengge Wu , Junsuo Zhao

Effective foundation modeling in remote sensing requires spatially aligned heterogeneous modalities coupled with semantically grounded supervision, yet such resources remain limited at scale. We present GeoMeld, a large-scale multimodal…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Maram Hasan , Md Aminur Hossain , Savitra Roy , Souparna Bhowmik , Ayush V. Patel , Mainak Singha , Subhasis Chaudhuri , Muhammad Haris Khan , Biplab Banerjee

With the current ubiquity of deep learning methods to solve computer vision and remote sensing specific tasks, the need for labelled data is growing constantly. However, in many cases, the annotation process can be long and tedious…

Computer Vision and Pattern Recognition · Computer Science 2023-06-19 Paul Berg , Minh-Tan Pham , Nicolas Courty
‹ Prev 1 2 3 10 Next ›