English
Related papers

Related papers: GeoMag: A Vision-Language Model for Pixel-level Fi…

200 papers

Recent advances in MLLMs are reframing segmentation from fixed-category prediction to instruction-grounded localization. While reasoning based segmentation has progressed rapidly in natural scenes, remote sensing lacks a generalizable…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Lifan Jiang , Yuhang Pei , oxi Wu , Yan Zhao , Tianrun Wu , Shulong Yu , Lihui Zhang , Deng Cai

Context modeling is critical for remote sensing image dense prediction tasks. Nowadays, the growing size of very-high-resolution (VHR) remote sensing images poses challenges in effectively modeling context. While transformer-based models…

Computer Vision and Pattern Recognition · Computer Science 2024-04-11 Sijie Zhao , Hao Chen , Xueliang Zhang , Pengfeng Xiao , Lei Bai , Wanli Ouyang

The advent of large vision-language models (LVLMs) represents a remarkable advance in the quest for artificial general intelligence. However, the model's effectiveness in both specialized and general tasks warrants further investigation.…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Yao Jiang , Xinyu Yan , Ge-Peng Ji , Keren Fu , Meijun Sun , Huan Xiong , Deng-Ping Fan , Fahad Shahbaz Khan

Gaze understanding unifies the detection of people, their gaze targets, and objects of interest into a single framework, offering critical insight into visual attention and intent estimation. Although prior research has modelled gaze cues…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Athul M. Mathew , Haithem Hermassi , Thariq Khalid , Arshad Ali Khan

Multi-source remote sensing enables complementary observation of ground objects, while cross-modal fine-grained object retrieval remains challenging, especially under unaligned optical and SAR conditions. Unlike conventional retrieval…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Tiantong Fang , Xiuwei Wang , Jing Xiao , Wujie Zhou , Liang Liao , Mi Wang

Vision Language Models (VLMs) are good at recognizing the global location of a photograph -- their geolocation prediction accuracy rivals the best human experts. But many VLMs are startlingly bad at \textit{explaining} which image evidence…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Mohit Talreja , Joshua Diao , Jim Thannikary James , Radu Casapu , Tejas Santanam , Ethan Mendes , Alan Ritter , Wei Xu , James Hays

Grounding natural language queries in graphical user interfaces (GUIs) presents a challenging task that requires models to comprehend diverse UI elements across various applications and systems, while also accurately predicting the spatial…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Zhecheng Li , Guoxian Song , Yiwei Wang , Zhen Xiong , Junsong Yuan , Yujun Cai

Image geo-localization is the task of predicting the specific location of an image and requires complex reasoning across visual, geographical, and cultural contexts. While prior Vision Language Models (VLMs) have the best accuracy at this…

Computation and Language · Computer Science 2025-02-21 Zheyuan Zhang , Runze Li , Tasnim Kabir , Jordan Boyd-Graber

The proliferation of remote sensing satellites has resulted in a massive amount of remote sensing images. However, due to human and material resource constraints, the vast majority of remote sensing images remain unlabeled. As a result, it…

Computer Vision and Pattern Recognition · Computer Science 2022-02-16 Wenyuan Li , Keyan Chen , Hao Chen , Zhenwei Shi

Vision language models (VLMs) can flexibly address various vision tasks through text interactions. Although successful in semantic understanding, state-of-the-art VLMs including GPT-5 still struggle in understanding 3D from 2D inputs. On…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Zhipeng Cai , Ching-Feng Yeh , Hu Xu , Zhuang Liu , Gregory Meyer , Xinjie Lei , Changsheng Zhao , Shang-Wen Li , Vikas Chandra , Yangyang Shi

In this work, we introduce the Geometry-Aware Large Reconstruction Model (GeoLRM), an approach which can predict high-quality assets with 512k Gaussians and 21 input images in only 11 GB GPU memory. Previous works neglect the inherent…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Chubin Zhang , Hongliang Song , Yi Wei , Yu Chen , Jiwen Lu , Yansong Tang

Large Language Model-based Vision-Language Models (LLM-based VLMs) have demonstrated impressive results in various vision-language understanding tasks. However, how well these VLMs can see image detail beyond the semantic level remains…

Computer Vision and Pattern Recognition · Computer Science 2024-08-08 Chenhui Gou , Abdulwahab Felemban , Faizan Farooq Khan , Deyao Zhu , Jianfei Cai , Hamid Rezatofighi , Mohamed Elhoseiny

Spatial intelligence requires visual representations that capture both semantic objects and geometric structure in the physical world. To support this, two major pre-training schemes are now widely used as foundation backbones:…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Haozhan Shen , Tiancheng Zhao , Kangjia Zhao , Jianwei Yin

Vision-Language-Action (VLA) models have emerged as a promising framework for enabling generalist robots capable of perceiving, reasoning, and acting in the real world. These models usually build upon pretrained Vision-Language Models…

Robotics · Computer Science 2025-11-25 Tao Lin , Gen Li , Yilei Zhong , Yanwen Zou , Yuxin Du , Jiting Liu , Encheng Gu , Bo Zhao

Radiology Report Generation (RRG) through Vision-Language Models (VLMs) promises to reduce documentation burden, improve reporting consistency, and accelerate clinical workflows. However, their clinical adoption remains limited by the lack…

Computer Vision and Pattern Recognition · Computer Science 2026-02-18 Marco Salmè , Federico Siciliano , Fabrizio Silvestri , Paolo Soda , Rosa Sicilia , Valerio Guarrasi

Automated textual description of remote sensing images is crucial for unlocking their full potential in diverse applications, from environmental monitoring to urban planning and disaster management. However, existing studies in remote…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Kaiyu Li , Zixuan Jiang , Xiangyong Cao , Jiayu Wang , Yuchen Xiao , Deyu Meng , Zhi Wang

Recently, large vision-language models (LVLMs) unleash powerful analysis capabilities for low Earth orbit (LEO) satellite Earth observation images in the data center. However, fast satellite motion, brief satellite-ground station (GS)…

Networking and Internet Architecture · Computer Science 2025-07-09 Yuxin Zhang , Jiahao Yang , Zhe Chen , Wenjun Zhu , Jin Zhao , Yue Gao

Fine-grained ship classification in remote sensing (RS-FGSC) poses a significant challenge due to the high similarity between classes and the limited availability of labeled data, limiting the effectiveness of traditional supervised…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Long Lan , Fengxiang Wang , Xiangtao Zheng , Zengmao Wang , Xinwang Liu

The recent development of vision language models (VLMs) has led to significant advances in visual-language integration through visual instruction tuning, and they have rapidly evolved in the field of remote sensing image understanding,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Kaixuan Lu

Real-world vision-language applications demand varying levels of perceptual granularity. However, most existing visual large language models (VLLMs), such as LLaVA, pre-assume a fixed resolution for downstream tasks, which leads to subpar…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Weiqing Luo , Zhen Tan , Yifan Li , Xinyu Zhao , Kwonjoon Lee , Behzad Dariush , Tianlong Chen