English
Related papers

Related papers: EarthBridge: A Solution for 4th Multi-modal Aerial…

200 papers

This paper explores the use of multi-conditional adversarial networks for SAR-to-EO image translation. Previous methods condition adversarial networks only on the input SAR. We show that incorporating multiple complementary modalities such…

Computer Vision and Pattern Recognition · Computer Science 2022-07-28 Armando Cabrera , Miriam Cha , Prafull Sharma , Michael Newey

Visual neural decoding seeks to reconstruct or infer perceived visual stimuli from brain activity patterns, providing critical insights into human cognition and enabling transformative applications in brain-computer interfaces and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Wenjiang Zhang , Sifeng Wang , Yuwei Su , Xinyu Li , Chen Zhang , Suyu Zhong

Vision-language foundation models achieve promising performance in natural image classification, yet their direct application to medical imaging is limited by severe domain shifts, resolution mismatches, and the multi-label nature of…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Yitong Li , Morteza Ghahremani , Christian Wachinger

Full 3D inversion of time-domain Airborne ElectroMagnetic (AEM) data requires specialists' expertise and a tremendous amount of computational resources, not readily available to everyone. Consequently, quasi-2D/3D inversion methods are…

Geophysics · Physics 2022-11-18 Wouter Deleersnyder , David Dudal , Thomas Hermans

The task of translating visible-to-infrared images (V2IR) is inherently challenging due to three main obstacles: 1) achieving semantic-aware translation, 2) managing the diverse wavelength spectrum in infrared imagery, and 3) the scarcity…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Lingyan Ran , Lidong Wang , Guangcong Wang , Peng Wang , Yanning Zhang

We present a novel framework for exemplar based image translation. Recent advanced methods for this task mainly focus on establishing cross-domain semantic correspondence, which sequentially dominates image generation in the manner of local…

Computer Vision and Pattern Recognition · Computer Science 2023-03-31 Chang Jiang , Fei Gao , Biao Ma , Yuhao Lin , Nannan Wang , Gang Xu

Deep learning has achieved some success in addressing the challenge of cloud removal in optical satellite images, by fusing with synthetic aperture radar (SAR) images. Recently, diffusion models have emerged as powerful tools for cloud…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Yuyang Hu , Suhas Lohit , Ulugbek S. Kamilov , Tim K. Marks

While the Earth observation community has witnessed a surge in high-impact foundation models and global Earth embedding datasets, a significant barrier remains in translating these academic assets into freely accessible tools. This tutorial…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Yijie Zheng , Weijie Wu , Bingyue Wu , Long Zhao , Guoqing Li , Mikolaj Czerkawski , Konstantin Klemmer

Multimodal transformer exhibits high capacity and flexibility to align image and text for visual grounding. However, the existing encoder-only grounding framework (e.g., TransVG) suffers from heavy computation due to the self-attention…

Computer Vision and Pattern Recognition · Computer Science 2023-10-27 Fengyuan Shi , Ruopeng Gao , Weilin Huang , Limin Wang

This paper presents a pilot study introducing a multimodal fusion framework for the detection and analysis of bridge defects, integrating Non-Destructive Evaluation (NDE) techniques with advanced image processing to enable precise…

Computer Vision and Pattern Recognition · Computer Science 2025-07-10 Ravi Datta Rachuri , Duoduo Liao , Samhita Sarikonda , Datha Vaishnavi Kondur

Image-to-image translation aims to learn the mapping between two visual domains. There are two main challenges for this task: 1) lack of aligned training pairs and 2) multiple possible outputs from a single input image. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2019-12-19 Hsin-Ying Lee , Hung-Yu Tseng , Qi Mao , Jia-Bin Huang , Yu-Ding Lu , Maneesh Singh , Ming-Hsuan Yang

Traversability estimation is critical for enabling robots to navigate across diverse terrains and environments. While recent self-supervised learning methods achieve promising results, they often fail to capture the characteristics of…

Robotics · Computer Science 2025-08-26 Zipeng Fang , Yanbo Wang , Lei Zhao , Weidong Chen

Open-vocabulary semantic segmentation (OVSS) involves assigning labels to each pixel in an image based on textual descriptions, leveraging world models like CLIP. However, they encounter significant challenges in cross-domain…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Aniruddh Sikdar , Aditya Gandhamal , Suresh Sundaram

Modern Earth observation (EO) increasingly leverages deep learning to harness the scale and diversity of satellite imagery across sensors and regions. While recent foundation models have demonstrated promising generalization across EO…

Aircraft-based surveying to collect airborne electromagnetic data is a key method to image large swaths of the Earth's surface in pursuit of better knowledge of aquifer systems. Despite many years of advancements, 3D inversion still poses…

Air-ground collaborative intelligence is becoming a key approach for next-generation urban intelligent transportation management, where aerial and ground systems work together on perception, communication, and decision-making. However, the…

Recent advances in multimodal large language models (LLMs) have led to significant progress in understanding, generation, and retrieval tasks. However, current solutions often treat these tasks in isolation or require training LLMs from…

Machine Learning · Computer Science 2025-09-24 Teng Xiao , Zuchao Li , Lefei Zhang

All-in-One Image Restoration (AiOIR) faces the fundamental challenge in reconciling conflicting optimization objectives across heterogeneous degradations. Existing methods are often constrained by coarse-grained control mechanisms or fixed…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Luwei Tu , Jiawei Wu , Xing Luo , Zhi Jin

Recent developments in application of deep learning models to acoustic Full Waveform Inversion (FWI) are marked by the use of diffusion models as prior distributions for Bayesian-like inference procedures. The advantage of these methods is…

Machine Learning · Computer Science 2025-06-19 A. S. Stankevich , I. B. Petrov

We present a method for trajectory planning for autonomous driving, learning image-based context embeddings that align with motion prediction frameworks and planning-based intention input. Within our method, a ViT encoder takes raw images…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Maitrayee Keskar , Mohan Trivedi , Ross Greer