中文
相关论文

相关论文: GeoPix: Multi-Modal Large Language Model for Pixel…

200 篇论文

Multi-modal Large Language Models (MLLMs) have shown remarkable capabilities across a wide range of vision-language tasks. However, due to the restricted input resolutions, MLLMs face significant challenges in precisely understanding and…

计算机视觉与模式识别 · 计算机科学 2025-10-27 Lu Zhang , Jiazuo Yu , Haomiao Xiong , Ping Hu , Yunzhi Zhuge , Huchuan Lu , You He

Remote Sensing Image Captioning (RSIC) is a cross-modal field bridging vision and language, aimed at automatically generating natural language descriptions of features and scenes in remote sensing imagery. Despite significant advances in…

计算机视觉与模式识别 · 计算机科学 2025-03-07 Qing Zhou , Tao Yang , Junyu Gao , Weiping Ni , Junzheng Wu , Qi Wang

Remote Sensing Image Captioning (RSIC) presents unique challenges and plays a critical role in applications. Traditional RSIC methods often struggle to produce rich and diverse descriptions. Recently, with advancements in VLMs, efforts have…

计算机视觉与模式识别 · 计算机科学 2025-02-12 Hui Lin , Danfeng Hong , Shuhang Ge , Chuyao Luo , Kai Jiang , Hao Jin , Congcong Wen

The application of Vision-language foundation models (VLFMs) to remote sensing (RS) imagery has garnered significant attention due to their superior capability in various downstream tasks. A key challenge lies in the scarcity of…

计算机视觉与模式识别 · 计算机科学 2025-07-23 Yiguo He , Junjie Zhu , Yiying Li , Xiaoyu Zhang , Chunping Qiu , Jun Wang , Qiangjuan Huang , Ke Yang

Recent advances in multimodal large language models (MLLMs) have expanded research in video understanding, primarily focusing on high-level tasks such as video captioning and question-answering. Meanwhile, a smaller body of work addresses…

计算机视觉与模式识别 · 计算机科学 2025-04-04 Ali Athar , Xueqing Deng , Liang-Chieh Chen

Recent advancements in Multimodal Large Language Models (MLLMs) have significantly enhanced performance on 2D visual tasks. However, improving their spatial intelligence remains a challenge. Existing 3D MLLMs always rely on additional 3D or…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Diankun Wu , Fangfu Liu , Yi-Hsin Hung , Yueqi Duan

Although multimodal large language models (MLLMs) have achieved promising results on a wide range of vision-language tasks, their ability to perceive and understand human faces is rarely explored. In this work, we comprehensively evaluate…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Haomiao Sun , Mingjie He , Tianheng Lian , Hu Han , Shiguang Shan

Large Multimodal Models (LMMs) extend Large Language Models to the vision domain. Initial LMMs used holistic images and text prompts to generate ungrounded textual responses. Recently, region-level LMMs have been used to generate visually…

Automated textual description of remote sensing images is crucial for unlocking their full potential in diverse applications, from environmental monitoring to urban planning and disaster management. However, existing studies in remote…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Kaiyu Li , Zixuan Jiang , Xiangyong Cao , Jiayu Wang , Yuchen Xiao , Deyu Meng , Zhi Wang

Achieving deep alignment between vision and language remains a central challenge for Multimodal Large Language Models (MLLMs). These models often fail to fully leverage visual input, defaulting to strong language priors. Our approach first…

计算机视觉与模式识别 · 计算机科学 2025-07-03 Aarti Ghatkesar , Ganesh Venkatesh

Spectral information has long been recognized as a critical cue in remote sensing observations. Although numerous vision-language models have been developed for pixel-level interpretation, spectral information remains underutilized,…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Dongchen Si , Di Wang , Erzhong Gao , Xiaolei Qin , Liu Zhao , Jing Zhang , Minqiang Xu , Jianbo Zhan , Jianshe Wang , Lin Liu , Bo Du , Liangpei Zhang

Methods based on Contrastive Language-Image Pre-training (CLIP) are nowadays extensively used in support of vision-and-language tasks involving remote sensing data, such as cross-modal retrieval. The adaptation of CLIP to this specific…

计算机视觉与模式识别 · 计算机科学 2024-11-01 João Daniel Silva , Joao Magalhaes , Devis Tuia , Bruno Martins

The rapid progress in Multimodal Large Language Models (MLLMs) has significantly advanced their ability to process and understand complex visual and textual information. However, the integration of multiple images and extensive textual…

计算机视觉与模式识别 · 计算机科学 2024-08-27 Yujie Lu , Xiujun Li , Tsu-Jui Fu , Miguel Eckstein , William Yang Wang

Multi-view understanding, the ability to reconcile visual information across diverse viewpoints for effective navigation, manipulation, and 3D scene comprehension, is a fundamental challenge in Multi-Modal Large Language Models (MLLMs) to…

计算机视觉与模式识别 · 计算机科学 2025-04-29 Chun-Hsiao Yeh , Chenyu Wang , Shengbang Tong , Ta-Ying Cheng , Ruoyu Wang , Tianzhe Chu , Yuexiang Zhai , Yubei Chen , Shenghua Gao , Yi Ma

The development of Large Vision-Language Models (LVLMs) is striving to catch up with the success of Large Language Models (LLMs), yet it faces more challenges to be resolved. Very recent works enable LVLMs to localize object-level visual…

计算机视觉与模式识别 · 计算机科学 2024-03-20 Zhipeng Huang , Zhizheng Zhang , Zheng-Jun Zha , Yan Lu , Baining Guo

Multimodal large language models (MLLMs) have altered the landscape of computer vision, obtaining impressive results across a wide range of tasks, especially in zero-shot settings. Unfortunately, their strong performance does not always…

计算机视觉与模式识别 · 计算机科学 2025-04-16 Darryl Hannan , John Cooper , Dylan White , Timothy Doster , Henry Kvinge , Yijing Watkins

Medical image segmentation is crucial for clinical diagnosis, yet existing models are limited by their reliance on explicit human instructions and lack the active reasoning capabilities to understand complex clinical questions. While recent…

计算机视觉与模式识别 · 计算机科学 2025-06-13 Yu Huang , Zelin Peng , Yichen Zhao , Piao Yang , Xiaokang Yang , Wei Shen

The multimodal fusion of images and scene captions has been extensively explored and applied in various fields. However, when dealing with complex remote sensing (RS) scenes, existing studies have predominantly concentrated on architectural…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Ziyi Wang , Xianping Ma , Ziyao Wang , Hongyang Zhang , Man On Pun

Remote sensing change detection (RSCD), a complex multi-image inference task, traditionally uses pixel-based operators or encoder-decoder networks that inadequately capture high-level semantics and are vulnerable to non-semantic…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Xingwei Ma , Shiyang Feng , Bo Zhang , Bin Wang

Large language models (LLMs) have recently been extended to the vision-language realm, obtaining impressive general multi-modal capabilities. However, the exploration of multi-modal large language models (MLLMs) for remote sensing (RS) data…

计算机视觉与模式识别 · 计算机科学 2024-01-19 Yang Zhan , Zhitong Xiong , Yuan Yuan