English
Related papers

Related papers: SOMA-1M: A Large-Scale SAR-Optical Multi-resolutio…

200 papers

We introduce SOMA, the Spatial Memory framework for Out-of-Vision Manipulation in Vision-Language-Action (VLA) models. Most existing VLAs implicitly assume that task-relevant objects are always visible, leading to brittle and reactive…

Robotics · Computer Science 2026-05-22 Pengteng Li , Weiyu Guo , He Zhang , Tiefu Cai , Xiao He , Yandong Guo , Hui Xiong

High-resolution satellite imagery is a key element for many Earth monitoring applications. Satellites such as Sentinel-2 feature characteristics that are favorable for super-resolution algorithms such as aliasing and band-misalignment.…

Computer Vision and Pattern Recognition · Computer Science 2023-04-18 Ngoc Long Nguyen , Jérémy Anger , Axel Davy , Pablo Arias , Gabriele Facciolo

Effective foundation modeling in remote sensing requires spatially aligned heterogeneous modalities coupled with semantically grounded supervision, yet such resources remain limited at scale. We present GeoMeld, a large-scale multimodal…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Maram Hasan , Md Aminur Hossain , Savitra Roy , Souparna Bhowmik , Ayush V. Patel , Mainak Singha , Subhasis Chaudhuri , Muhammad Haris Khan , Biplab Banerjee

Due to its cloud-penetrating capability and independence from solar illumination, satellite Synthetic Aperture Radar (SAR) is the preferred data source for large-scale flood mapping, providing global coverage and including various land…

Computer Vision and Pattern Recognition · Computer Science 2024-06-07 Jie Zhao , Zhitong Xiong , Xiao Xiang Zhu

Image captioning has become an important task in computer vision, enabling models to generate natural language descriptions of visual content. While several datasets exist for natural images and high-resolution optical remote sensing…

Computer Vision and Pattern Recognition · Computer Science 2026-05-06 Lucrezia Tosato , Gianluca Lombardi , Ronny Hansch

Object extraction and segmentation from remote sensing (RS) images is a critical yet challenging task in urban environment monitoring. Urban morphology is inherently complex, with irregular objects of diverse shapes and varying scales.…

Computer Vision and Pattern Recognition · Computer Science 2025-02-24 Chenyu Li , Danfeng Hong , Bing Zhang , Yuxuan Li , Gustau Camps-Valls , Xiao Xiang Zhu , Jocelyn Chanussot

Geometric information in the normalized digital surface models (nDSM) is highly correlated with the semantic class of the land cover. Exploiting two modalities (RGB and nDSM (height)) jointly has great potential to improve the segmentation…

Computer Vision and Pattern Recognition · Computer Science 2023-05-25 Zhitong Xiong , Sining Chen , Yi Wang , Lichao Mou , Xiao Xiang Zhu

Drone-based multi-object tracking is essential yet highly challenging due to small targets, severe occlusions, and cluttered backgrounds. Existing RGB-based tracking algorithms heavily depend on spatial appearance cues such as color and…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Tianhao Li , Tingfa Xu , Ying Wang , Haolin Qin , Xu Lin , Jianan Li

We introduce MOMO, the first multi-sensor foundation model for Mars remote sensing. MOMO uses model merge to integrate representations learned independently from three key Martian sensors (HiRISE, CTX, and THEMIS), spanning resolutions from…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Mirali Purohit , Bimal Gajera , Irish Mehta , Bhanu Tokas , Jacob Adler , Steven Lu , Scott Dickenshied , Serina Diniega , Brian Bue , Umaa Rebbapragada , Hannah Kerner

With the proliferation of low altitude unmanned aerial vehicles (UAVs), visual multi-object tracking is becoming a critical security technology, demanding significant robustness even in complex environmental conditions. However, tracking…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Tianyang Xu , Jinjie Gu , Xuefeng Zhu , XiaoJun Wu , Josef Kittler

Synthesizing missing modalities in multi-modal magnetic resonance imaging (MRI) is vital for ensuring diagnostic completeness, particularly when full acquisitions are infeasible due to time constraints, motion artifacts, and patient…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Yue Zhang , Zhizheng Zhuo , Siyao Xu , Shan Lv , Zhaoxi Liu , Jun Qiu , Qiuli Wang , Yaou Liu , S. Kevin Zhou

Research on Multi-modal Large Language Models (MLLMs) towards the multi-image cross-modal instruction has received increasing attention and made significant progress, particularly in scenarios involving closely resembling images (e.g.,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-26 Tao Wu , Mengze Li , Jingyuan Chen , Wei Ji , Wang Lin , Jinyang Gao , Kun Kuang , Zhou Zhao , Fei Wu

Instruction-driven segmentation in remote sensing generates masks from guidance, offering great potential for accessible and generalizable applications. However, existing methods suffer from fragmented task formulations and limited…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Shuo Ni , Di Wang , He Chen , Haonan Guo , Ning Zhang , Jing Zhang

With the advancement of drone technology, the volume of video data increases rapidly, creating an urgent need for efficient semantic retrieval. We are the first to systematically propose and study the drone video-text retrieval (DVTR) task.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-20 Jinghao Huang , Yaxiong Chen , Ganchao Liu

Multimodal fusion of remote sensing images serves as a core technology for overcoming the limitations of single-source data and improving the accuracy of surface information extraction, which exhibits significant application value in fields…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Siyu Zhang , Lianlei Shan , Runhe Qiu

Recently, the diffusion-based generative paradigm has achieved impressive general image generation capabilities with text prompts due to its accurate distribution modeling and stable training process. However, generating diverse remote…

Image and Video Processing · Electrical Eng. & Systems 2024-10-31 Jialin Luo , Yuanzhi Wang , Ziqi Gu , Yide Qiu , Shuaizhen Yao , Fuyun Wang , Chunyan Xu , Wenhua Zhang , Dan Wang , Zhen Cui

Multimodal fusion has become a key enabler for UAV-based object detection, as each modality provides complementary cues for robust feature extraction. However, due to significant differences in resolution, field of view, and sensing…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Kangcheng Bin , Chen Chen , Ting Hu , Jiahao Qi , Ping Zhong

Multi-modal remote sensing images are vital for Earth observation, yet complete paired observations are often scarce in practice. Existing generative methods commonly address this problem through isolated pairwise modality translation, but…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Zhiping Yu , Chenyang Liu , Jinqi Cao , Qinzhe Yang , Siwei Yu , Zhengxia Zou , Zhenwei Shi

Multimodal change detection (MMCD) identifies changed areas in multimodal remote sensing (RS) data, demonstrating significant application value in land use monitoring, disaster assessment, and urban sustainable development. However,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Xuanguang Liu , Lei Ding , Yujie Li , Chenguang Dai , Zhenchao Zhang , Mengmeng Li , Ziyi Yang , Yifan Sun , Yongqi Sun , Hanyun Wang

Current multimodal and multitask foundation models like 4M or UnifiedIO show promising results, but in practice their out-of-the-box abilities to accept diverse inputs and perform diverse tasks are limited by the (usually rather small)…

Computer Vision and Pattern Recognition · Computer Science 2024-06-17 Roman Bachmann , Oğuzhan Fatih Kar , David Mizrahi , Ali Garjani , Mingfei Gao , David Griffiths , Jiaming Hu , Afshin Dehghan , Amir Zamir
‹ Prev 1 4 5 6 7 8 10 Next ›