English
Related papers

Related papers: MMM-RS: A Multi-modal, Multi-GSD, Multi-scene Remo…

200 papers

Neural Radiance Fields (NeRF) have shown impressive performances in the rendering of 3D scenes from arbitrary viewpoints. While RGB images are widely preferred for training volume rendering models, the interest in other radiance modalities…

Graphics · Computer Science 2025-03-26 Federico Lincetto , Gianluca Agresti , Mattia Rossi , Pietro Zanuttigh

Recent advances in Retrieval-Augmented Generation (RAG) have significantly improved response accuracy and relevance by incorporating external knowledge into Large Language Models (LLMs). However, existing RAG methods primarily focus on…

Machine Learning · Computer Science 2025-04-22 Qinhan Yu , Zhiyou Xiao , Binghui Li , Zhengren Wang , Chong Chen , Wentao Zhang

Recent advances in tuning-free personalized image generation based on diffusion models are impressive. However, to improve subject fidelity, existing methods either retrain the diffusion model or infuse it with dense visual embeddings, both…

Computer Vision and Pattern Recognition · Computer Science 2024-03-25 Zhichao Wei , Qingkun Su , Long Qin , Weizhi Wang

Recent advancements in subject-driven image generation have made significant strides. However, current methods still fall short in diverse application scenarios, as they require test-time tuning and cannot accept interleaved multi-image and…

Computer Vision and Pattern Recognition · Computer Science 2024-04-29 Xichen Pan , Li Dong , Shaohan Huang , Zhiliang Peng , Wenhu Chen , Furu Wei

We review research on generating visual data from text from the angle of "cross-modal generation." This point of view allows us to draw parallels between various methods geared towards working on input text and producing visual output,…

Computer Vision and Pattern Recognition · Computer Science 2024-01-23 Maciej Żelaszczyk , Jacek Mańdziuk

This research paper proposes a Latent Diffusion Model for 3D (LDM3D) that generates both image and depth map data from a given text prompt, allowing users to generate RGBD images from text prompts. The LDM3D model is fine-tuned on a dataset…

Computer Vision and Pattern Recognition · Computer Science 2023-05-23 Gabriela Ben Melech Stan , Diana Wofk , Scottie Fox , Alex Redden , Will Saxton , Jean Yu , Estelle Aflalo , Shao-Yen Tseng , Fabio Nonato , Matthias Muller , Vasudev Lal

Artificial intelligence generative content (AIGC) has significantly impacted image generation in the field of remote sensing. However, the equally important area of remote sensing image (RSI) editing has not received sufficient attention.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Fangzhou Han , Lingyu Si , Zhizhuo Jiang , Hongwei Dong , Lamei Zhang , Yu Liu , Hao Chen , Bo Du

Few-shot object detection (FSOD) aims to detect novel instances with only a limited number of labeled training samples, presenting a challenge that is particularly prominent in numerous remote sensing applications such as endangered species…

Image and Video Processing · Electrical Eng. & Systems 2025-11-25 Yanxing Liu , Jiancheng Pan , Jianwei Yang , Tiancheng Chen , Peiling Zhou , Bingchen Zhang

This work introduces composed image retrieval to remote sensing. It allows to query a large image archive by image examples alternated by a textual description, enriching the descriptive power over unimodal queries, either visual or…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Bill Psomas , Ioannis Kakogeorgiou , Nikos Efthymiadis , Giorgos Tolias , Ondrej Chum , Yannis Avrithis , Konstantinos Karantzalos

Context modeling is critical for remote sensing image dense prediction tasks. Nowadays, the growing size of very-high-resolution (VHR) remote sensing images poses challenges in effectively modeling context. While transformer-based models…

Computer Vision and Pattern Recognition · Computer Science 2024-04-11 Sijie Zhao , Hao Chen , Xueliang Zhang , Pengfeng Xiao , Lei Bai , Wanli Ouyang

Vision Language Foundation Models based on CLIP architecture for remote sensing primarily rely on short text captions, which often result in incomplete semantic representations. Although longer captions convey richer information, existing…

Computer Vision and Pattern Recognition · Computer Science 2025-10-30 Weizhi Chen , Yupeng Deng , Jin Wei , Jingbo Chen , Jiansheng Chen , Yuman Feng , Zhihao Xi , Diyou Liu , Kai Li , Yu Meng

Recent advances in diffusion models have led to a quantum leap in the quality of generative visual content. However, quantification of realism of the content is still challenging. Existing evaluation metrics, such as Inception Score and…

Computer Vision and Pattern Recognition · Computer Science 2023-09-27 Yunzhuo Chen , Naveed Akhtar , Nur Al Hasan Haldar , Ajmal Mian

Multimodal remote sensing image (MRSI) matching is pivotal for cross-modal fusion, localization, and object detection, but it faces severe challenges due to geometric, radiometric, and viewpoint discrepancies across imaging modalities.…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Peihao Wu , Yongxiang Yao , Wenfei Zhang , Dong Wei , Yi Wan , Yansheng Li , Yongjun Zhang

With the rapid progress of controllable generation, training data synthesis has become a promising way to expand labeled datasets and alleviate manual annotation in remote sensing (RS). However, the complexity of semantic mask control and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Yunkai Yang , Yudong Zhang , Kunquan Zhang , Jinxiao Zhang , Xinying Chen , Haohuan Fu , Runmin Dong

In this paper, we introduce the task of visual grounding for remote sensing data (RSVG). RSVG aims to localize the referred objects in remote sensing (RS) images with the guidance of natural language. To retrieve rich information from RS…

Computer Vision and Pattern Recognition · Computer Science 2023-05-03 Yang Zhan , Zhitong Xiong , Yuan Yuan

Masked generative models (MGMs) have shown impressive generative ability while providing an order of magnitude efficient sampling steps compared to continuous diffusion models. However, MGMs still underperform in image synthesis compared to…

Computer Vision and Pattern Recognition · Computer Science 2024-10-18 Jiwan Hur , Dong-Jae Lee , Gyojin Han , Jaehyun Choi , Yunho Jeon , Junmo Kim

Image fusion in Remote Sensing (RS) has been a consistent demand due to its ability to turn raw images of different resolutions, sources, and modalities into accurate, complete, and spatio-temporally coherent images. It greatly facilitates…

Computer Vision and Pattern Recognition · Computer Science 2024-01-18 Hessah Albanwan , Rongjun Qin , Yang Tang

Dataset distillation compresses large training sets into compact synthetic datasets while preserving downstream performance. As modern systems increasingly operate on paired vision-language inputs, multimodal distillation must preserve…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Jongoh Jeong , Hoyong Kwon , Minseok Kim , Kuk-Jin Yoon

While deep learning has recently achieved great success on multi-view stereo (MVS), limited training data makes the trained model hard to be generalized to unseen scenarios. Compared with other computer vision tasks, it is rather difficult…

Computer Vision and Pattern Recognition · Computer Science 2020-04-14 Yao Yao , Zixin Luo , Shiwei Li , Jingyang Zhang , Yufan Ren , Lei Zhou , Tian Fang , Long Quan

Recent endeavors in Multimodal Large Language Models (MLLMs) aim to unify visual comprehension and generation by combining LLM and diffusion models, the state-of-the-art in each task, respectively. Existing approaches rely on spatial visual…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Kaihang Pan , Wang Lin , Zhongqi Yue , Tenglong Ao , Liyu Jia , Wei Zhao , Juncheng Li , Siliang Tang , Hanwang Zhang
‹ Prev 1 3 4 5 6 7 10 Next ›