English
Related papers

Related papers: GeoChat: Grounded Large Vision-Language Model for …

200 papers

Current large vision-language models (LVLMs) typically rely on text-only reasoning based on a single-pass visual encoding, which often leads to loss of fine-grained visual information. Recently the proposal of ''thinking with images''…

Computer Vision and Pattern Recognition · Computer Science 2026-02-13 Junfei Wu , Jian Guan , Qiang Liu , Shu Wu , Liang Wang , Wei Wu , Tieniu Tan

Natural-language-to-visualization (NL2VIS) systems based on large language models (LLMs) have substantially improved the accessibility of data visualization. However, their further adoption is hindered by two coupled challenges: (i) the…

Human-Computer Interaction · Computer Science 2026-01-23 Marko Hostnik , Rauf Kurbanov , Yaroslav Sokolov , Artem Trofimov

Cross-view geo-localisation identifies coarse geographical position of an automated vehicle by matching a ground-level image to a geo-tagged satellite image from a database. Despite the advancements in Cross-view geo-localisation,…

Computer Vision and Pattern Recognition · Computer Science 2025-05-21 Barkin Dagda , Muhammad Awais , Saber Fallah

Recently, the flourishing large language models(LLM), especially ChatGPT, have shown exceptional performance in language understanding, reasoning, and interaction, attracting users and researchers from multiple fields and domains. Although…

Computer Vision and Pattern Recognition · Computer Science 2024-01-18 Haonan Guo , Xin Su , Chen Wu , Bo Du , Liangpei Zhang , Deren Li

Geo-localization is the task of identifying the location of an image using visual cues alone. It has beneficial applications, such as improving disaster response, enhancing navigation, and geography education. Recently, Vision-Language…

Computer Vision and Pattern Recognition · Computer Science 2025-08-28 Oliver Grainge , Sania Waheed , Jack Stilgoe , Michael Milford , Shoaib Ehsan

Visual Question Answering (VQA) in remote sensing (RS) is pivotal for interpreting Earth observation data. However, existing RS VQA datasets are constrained by limitations in annotation richness, question diversity, and the assessment of…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Xing Zi , Jinghao Xiao , Yunxiao Shi , Xian Tao , Jun Li , Ali Braytee , Mukesh Prasad

Vision-Language Models (VLMs) have demonstrated great potential in interpreting remote sensing (RS) images through language-guided semantic. However, the effectiveness of these VLMs critically depends on high-quality image-text training…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Dilxat Muhtar , Enzhuo Zhang , Zhenshi Li , Feng Gu , Yanglangxing He , Pengfeng Xiao , Xueliang Zhang

Earth vision has achieved milestones in geospatial object recognition but lacks exploration in object-relational reasoning, limiting comprehensive scene understanding. To address this, a progressive Earth vision-language understanding and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Junjue Wang , Yanfei Zhong , Zihang Chen , Zhuo Zheng , Ailong Ma , Liangpei Zhang

Grounding language to the visual observations of a navigating agent can be performed using off-the-shelf visual-language models pretrained on Internet-scale data (e.g., image captions). While this is useful for matching images to natural…

Robotics · Computer Science 2023-03-09 Chenguang Huang , Oier Mees , Andy Zeng , Wolfram Burgard

Multimodal large language models (MLLMs) have altered the landscape of computer vision, obtaining impressive results across a wide range of tasks, especially in zero-shot settings. Unfortunately, their strong performance does not always…

Computer Vision and Pattern Recognition · Computer Science 2025-04-16 Darryl Hannan , John Cooper , Dylan White , Timothy Doster , Henry Kvinge , Yijing Watkins

Ensuring accessible pedestrian navigation requires reasoning about both semantic and spatial aspects of complex urban scenes, a challenge that existing Large Vision-Language Models (LVLMs) struggle to meet. Although these models can…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Rafi Ibn Sultan , Hui Zhu , Xiangyu Zhou , Chengyin Li , Prashant Khanduri , Marco Brocanelli , Dongxiao Zhu

In human reading and communication, individuals tend to engage in geospatial reasoning, which involves recognizing geographic entities and making informed inferences about their interrelationships. To mimic such cognitive process, current…

Computation and Language · Computer Science 2024-08-22 Yibo Yan , Joey Lee

Automated textual description of remote sensing images is crucial for unlocking their full potential in diverse applications, from environmental monitoring to urban planning and disaster management. However, existing studies in remote…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Kaiyu Li , Zixuan Jiang , Xiangyong Cao , Jiayu Wang , Yuchen Xiao , Deyu Meng , Zhi Wang

We introduce a new benchmark designed to advance the development of general-purpose, large-scale vision-language models for remote sensing images. Although several vision-language datasets in remote sensing have been proposed to pursue this…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Xiang Li , Jian Ding , Mohamed Elhoseiny

Abundant, well-annotated multimodal data in remote sensing are pivotal for aligning complex visual remote sensing (RS) scenes with human language, enabling the development of specialized vision language models across diverse RS…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Junyao Ge , Xu Zhang , Yang Zheng , Kaitai Guo , Jimin Liang

Large Vision--Language Models (LVLMs) hold great promise for advancing optical remote sensing (RS) analysis, yet existing reasoning segmentation frameworks couple linguistic reasoning and pixel prediction through end-to-end supervised…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Xu Zhang , Junyao Ge , Yang Zheng , Kaitai Guo , Jimin Liang

We introduce Groma, a Multimodal Large Language Model (MLLM) with grounded and fine-grained visual perception ability. Beyond holistic image understanding, Groma is adept at region-level tasks such as region captioning and visual grounding.…

Computer Vision and Pattern Recognition · Computer Science 2024-04-22 Chuofan Ma , Yi Jiang , Jiannan Wu , Zehuan Yuan , Xiaojuan Qi

Classifying geospatial imagery remains a major bottleneck for applications such as disaster response and land-use monitoring-particularly in regions where annotated data is scarce or unavailable. Existing tools (e.g., RS-CLIP) that claim…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Gilles Quentin Hacheme , Girmaw Abebe Tadesse , Caleb Robinson , Akram Zaytar , Rahul Dodhia , Juan M. Lavista Ferres

Segmentation models can recognize a pre-defined set of objects in images. However, models that can reason over complex user queries that implicitly refer to multiple objects of interest are still in their infancy. Recent advances in…

Artificial Intelligence · Computer Science 2025-05-06 Jerome Quenum , Wen-Han Hsieh , Tsung-Han Wu , Ritwik Gupta , Trevor Darrell , David M. Chan

In the field of multimodal chain-of-thought (CoT) reasoning, existing approaches predominantly rely on reasoning on pure language space, which inherently suffers from language bias and is largely confined to math or science domains. This…

Computer Vision and Pattern Recognition · Computer Science 2026-05-04 Jiacong Wang , Zijian Kang , Haochen Wang , Haiyong Jiang , Jiawen Li , Bohong Wu , Ya Wang , Jiao Ran , Xiao Liang , Chao Feng , Jun Xiao
‹ Prev 1 3 4 5 6 7 10 Next ›