English
Related papers

Related papers: CDChat: A Large Multimodal Model for Remote Sensin…

200 papers

Recent advances in multimodal large language models (MLLMs) have demonstrated impressive results in various visual tasks. However, in remote sensing (RS), high resolution and small proportion of objects pose challenges to existing MLLMs,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Hongxiang Jiang , Jihao Yin , Qixiong Wang , Jiaqi Feng , Guo Chen

Large multimodal models (LMM) have recently shown encouraging progress with visual instruction tuning. In this note, we show that the fully-connected vision-language cross-modal connector in LLaVA is surprisingly powerful and…

Computer Vision and Pattern Recognition · Computer Science 2024-05-17 Haotian Liu , Chunyuan Li , Yuheng Li , Yong Jae Lee

Multimodal large language models (MLLMs) have made rapid progress in recent years, yet continue to struggle with low-level visual perception (LLVP) -- particularly the ability to accurately describe the geometric details of an image. This…

Computer Vision and Pattern Recognition · Computer Science 2024-12-13 Jiarui Zhang , Ollie Liu , Tianyu Yu , Jinyi Hu , Willie Neiswanger

Recent advancements in large vision-language models (LVLMs), such as GPT4-V and LLaVA, have been substantial. LLaVA's modular architecture, in particular, offers a blend of simplicity and efficiency. Recent works mainly focus on introducing…

Computer Vision and Pattern Recognition · Computer Science 2024-05-21 Yuan Liu , Le Tian , Xiao Zhou , Jie Zhou

We systematically investigate lightweight strategies to adapt large language models (LLMs) for the task of radiology report summarization (RRS). Specifically, we focus on domain adaptation via pretraining (on natural language, biomedical…

This work investigates the use of large language models (LLMs) for tasks in smart cities. The core idea is to leverage remote sensing imagery to characterize the built environment, including design suggestions, constructability assessment,…

Computation and Language · Computer Science 2026-05-12 Dongdong Wang , Deepak Balakrishnan , Ravi Srinivasan , Shenhao Wang

Multimodal Large Language Models (MLLMs) have made significant progress in tasks such as image captioning and question answering. However, while these models can generate realistic captions, they often struggle with providing precise…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Chun-Peng Chang , Alain Pagani , Didier Stricker

Few-shot learning has been studied to adapt models to tasks with very few samples. It holds profound significance, particularly in clinical tasks, due to the high annotation cost of medical images. Several works have explored few-shot…

Computer Vision and Pattern Recognition · Computer Science 2024-02-06 Kaipeng Zheng , Weiran Huang , Lichao Sun

The remote sensing community has recently seen the emergence of methods based on Large Vision and Language Models (LVLMs) that can address multiple tasks at the intersection of computer vision and natural language processing. To fully…

Computer Vision and Pattern Recognition · Computer Science 2025-12-18 João Daniel Silva , Joao Magalhaes , Devis Tuia , Bruno Martins

Large language models (LLMs) and large multimodal models (LMMs) have significantly impacted the AI community, industry, and various economic sectors. In journalism, integrating AI poses unique challenges and opportunities, particularly in…

Computation and Language · Computer Science 2024-08-09 Aliki Anagnostopoulou , Thiago Gouvea , Daniel Sonntag

Large language models (LLMs) have recently demonstrated their potential in clinical applications, providing valuable medical knowledge and advice. For example, a large dialog LLM like ChatGPT has successfully passed part of the US medical…

Computer Vision and Pattern Recognition · Computer Science 2023-02-15 Sheng Wang , Zihao Zhao , Xi Ouyang , Qian Wang , Dinggang Shen

Recent advances in large multimodal models (LMMs) have recognized fine-grained grounding as an imperative factor of visual understanding and dialogue. However, the benefits of such representation in LMMs are limited to the natural image…

Computer Vision and Pattern Recognition · Computer Science 2025-01-24 Akashah Shabbir , Mohammed Zumri , Mohammed Bennamoun , Fahad S. Khan , Salman Khan

Multimodal large language models (MLLMs) excel at 2D visual understanding but remain limited in their ability to reason about 3D space. In this work, we leverage large-scale high-quality 3D scene data with open-set annotations to introduce…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Erik Daxberger , Nina Wenzel , David Griffiths , Haiming Gang , Justin Lazarow , Gefen Kohavi , Kai Kang , Marcin Eichner , Yinfei Yang , Afshin Dehghan , Peter Grasch

While numerous recent benchmarks focus on evaluating generic Vision-Language Models (VLMs), they do not effectively address the specific challenges of geospatial applications. Generic VLM benchmarks are not designed to handle the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Muhammad Sohail Danish , Muhammad Akhtar Munir , Syed Roshaan Ali Shah , Kartik Kuckreja , Fahad Shahbaz Khan , Paolo Fraccaro , Alexandre Lacoste , Salman Khan

Existing visual instruction tuning methods typically prompt large language models with textual descriptions to generate instruction-following data. Despite the promising performance achieved, these descriptions are derived from image…

Computer Vision and Pattern Recognition · Computer Science 2023-11-30 Junke Wang , Lingchen Meng , Zejia Weng , Bo He , Zuxuan Wu , Yu-Gang Jiang

Large Multimodal Models (LMMs) have achieved strong performance across a range of vision and language tasks. However, their spatial reasoning capabilities are under-investigated. In this paper, we construct a novel VQA dataset, Spatial-MM,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Fatemeh Shiri , Xiao-Yu Guo , Mona Golestan Far , Xin Yu , Gholamreza Haffari , Yuan-Fang Li

Understanding the deep semantics of images is essential in the era dominated by social media. However, current research works primarily on the superficial description of images, revealing a notable deficiency in the systematic investigation…

Computation and Language · Computer Science 2024-06-21 Yixin Yang , Zheng Li , Qingxiu Dong , Heming Xia , Zhifang Sui

Recent progress in Multimodal Large Language Models (MLLMs) has highlighted the critical roles of both the visual backbone and the underlying language model. While prior work has primarily focused on scaling these components to billions of…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 Federico Cocchi , Nicholas Moratelli , Davide Caffagni , Sara Sarto , Lorenzo Baraldi , Marcella Cornia , Rita Cucchiara

Instruction finetuning is a popular paradigm to align large language models (LLM) with human intent. Despite its popularity, this idea is less explored in improving LLMs to align existing foundation models with scientific disciplines,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Sameera Horawalavithana , Sai Munikoti , Ian Stewart , Henry Kvinge , Karl Pazdernik

Large vision-language models (LVLMs) have been regarded as a breakthrough advance in an astoundingly variety of tasks, from content generation to virtual assistants and multimodal search or retrieval. However, for many of these…

Computer Vision and Pattern Recognition · Computer Science 2025-02-03 Kailash Hambarde , Pranita Samale , Hugo Proença