English
Related papers

Related papers: GEOBIND: Binding Text, Image, and Audio through Sa…

200 papers

Building detection from satellite multispectral imagery data is being a fundamental but a challenging problem mainly because it requires correct recovery of building footprints from high-resolution images. In this work, we propose a deep…

Computer Vision and Pattern Recognition · Computer Science 2020-10-12 Geesara Prathap , Ilya Afanasyev

We address the problem of estimating depth with multi modal audio visual data. Inspired by the ability of animals, such as bats and dolphins, to infer distance of objects with echolocation, some recent methods have utilized echoes for depth…

Computer Vision and Pattern Recognition · Computer Science 2021-04-06 Kranti Kumar Parida , Siddharth Srivastava , Gaurav Sharma

Multimodal sensing systems are increasingly prevalent in various real-world applications. Most existing multimodal learning approaches heavily rely on training with a large amount of synchronized, complete multimodal data. However, such a…

Machine Learning · Computer Science 2025-03-06 Xiaomin Ouyang , Jason Wu , Tomoyoshi Kimura , Yihan Lin , Gunjan Verma , Tarek Abdelzaher , Mani Srivastava

Recent urbanization has coincided with the enrichment of geotagged data, such as street view and point-of-interest (POI). Region embedding enhanced by the richer data modalities has enabled researchers and city administrators to understand…

Machine Learning · Computer Science 2021-05-07 Tianyuan Huang , Zhecheng Wang , Hao Sheng , Andrew Y. Ng , Ram Rajagopal

Multilingual (or cross-lingual) embeddings represent several languages in a unique vector space. Using a common embedding space enables for a shared semantic between words from different languages. In this paper, we propose to embed images…

Computer Vision and Pattern Recognition · Computer Science 2019-05-15 Maxime Portaz , Hicham Randrianarivo , Adrien Nivaggioli , Estelle Maudet , Christophe Servan , Sylvain Peyronnet

A satellite image is a remotely sensed image data, where each pixel represents a specific location on earth. The pixel value recorded is the reflection radiation from the earth's surface at that location. Multispectral images are those that…

Computer Vision and Pattern Recognition · Computer Science 2022-03-22 Purbarag Pathak Choudhury , Ujjal Kr Dutta , Dhruba Kr Bhattacharyya

Publicly available satellite imagery can be an ubiquitous, cheap, and powerful tool for vehicle localisation when a prior sensor map is unavailable. However, satellite images are not directly comparable to data from ground range sensors…

Robotics · Computer Science 2020-09-24 Tim Y. Tang , Daniele De Martini , Shangzhe Wu , Paul Newman

Seismic imaging is the numerical process of creating a volumetric representation of the subsurface geological structures from elastic waves recorded at the surface of the Earth. As such, it is widely utilized in the energy and construction…

Geophysics · Physics 2024-11-05 Juan Romero , Wolfgang Heidrich , Nick Luiken , Matteo Ravasi

In this paper, we attempt to address the challenging problem of counting built-structures in the satellite imagery. Building density is a more accurate estimate of the population density, urban area expansion and its impact on the…

Computer Vision and Pattern Recognition · Computer Science 2019-04-02 Anza Shakeel , Waqas Sultani , Mohsen Ali

In training machine learning models for land cover semantic segmentation there is a stark contrast between the availability of satellite imagery to be used as inputs and ground truth data to enable supervised learning. While thousands of…

Computer Vision and Pattern Recognition · Computer Science 2022-03-14 Michail Tarasiou , Stefanos Zafeiriou

We present three multi-scale similarity learning architectures, or DeepSim networks. These models learn pixel-level matching with a contrastive loss and are agnostic to the geometry of the considered scene. We establish a middle ground…

Computer Vision and Pattern Recognition · Computer Science 2023-04-18 Mohamed Ali Chebbi , Ewelina Rupnik , Marc Pierrot-Deseilligny , Paul Lopes

Recent work has shown that deep learning models can be used to classify land-use data from geospatial satellite imagery. We show that when these deep learning models are trained on data from specific continents/seasons, there is a high…

Computer Vision and Pattern Recognition · Computer Science 2021-06-21 Lucas Hu , Caleb Robinson , Bistra Dilkina

Remote sensing (RS) visual grounding aims to use natural language expression to locate specific objects (in the form of the bounding box or segmentation mask) in RS images, enhancing human interaction with intelligent RS interpretation…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Yue Zhou , Mengcheng Lan , Xiang Li , Litong Feng , Yiping Ke , Xue Jiang , Qingyun Li , Xue Yang , Wayne Zhang

The large variation of viewpoint and irrelevant content around the target always hinder accurate image retrieval and its subsequent tasks. In this paper, we investigate an extremely challenging task: given a ground-view image of a landmark,…

Computer Vision and Pattern Recognition · Computer Science 2022-05-24 Zelong Zeng , Zheng Wang , Fan Yang , Shin'ichi Satoh

Nowadays, as cameras are rapidly adopted in our daily routine, images of documents are becoming both abundant and prevalent. Unlike natural images that capture physical objects, document-images contain a significant amount of text with…

Computer Vision and Pattern Recognition · Computer Science 2021-03-19 Or Perel , Oron Anschel , Omri Ben-Eliezer , Shai Mazor , Hadar Averbuch-Elor

We present a deep learning approach for learning the joint semantic embeddings of images and captions in a Euclidean space, such that the semantic similarity is approximated by the L2 distances in the embedding space. For that, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2022-10-11 Noam Malali , Yosi Keller

Despite recent progress in Multi-Modal Large Language Models (MLLMs), it remains challenging to integrate diverse tasks ranging from pixel-level perception to high-fidelity generation. Existing approaches often suffer from either restricted…

Computation and Language · Computer Science 2026-01-29 Bin Zhu , Munan Ning , Peng Jin , Bin Lin , Jinfa Huang , Qi Song , Junwu Zhang , Zhenyu Tang , Mingjun Pan , Li Yuan

We describe a novel metric-based learning approach that introduces a multimodal framework and uses deep audio and geophone encoders in siamese configuration to design an adaptable and lightweight supervised model. This framework eliminates…

Sound · Computer Science 2021-11-16 Muhammad Shakeel , Katsutoshi Itoyama , Kenji Nishida , Kazuhiro Nakadai

This letter proposes a physics-aware multi-modal contrastive learning framework designed to transform complex seismic wavefields into human-readable physical representations. Traditional data-driven inversion methods often focus on…

Geophysics · Physics 2026-01-22 Chaohua Liang , Jun Matsushima

A soundscape is defined by the acoustic environment a person perceives at a location. In this work, we propose a framework for mapping soundscapes across the Earth. Since soundscapes involve sound distributions that span varying spatial…