中文
相关论文

相关论文: Explaining the Unseen: Multimodal Vision-Language …

200 篇论文

Detecting disasters in underground mining, such as explosions and structural damage, has been a persistent challenge over the years. This problem is compounded for first responders, who often have no clear information about the extent or…

计算机视觉与模式识别 · 计算机科学 2024-11-21 Mizanur Rahman Jewel , Mohamed Elmahallawy , Sanjay Madria , Samuel Frimpong

General-purpose vision-language models (VLMs) such as LLaVA and QwenVL produce descriptions of disaster imagery that lack domain-specific vocabulary and actionable detail. We propose the Vision-Language Caption Enhancer (VLCE), a framework…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Md. Mahfuzur Rahman , Kishor Datta Gupta , Marufa Kamal , Fahad Rahman , Sunzida Siddique , Ahmed Rafi Hasan , Mohd Ariful Haque , Roy George

Large vision-language models (VLMs) have made great achievements in Earth vision. However, complex disaster scenes with diverse disaster types, geographic regions, and satellite sensors have posed new challenges for VLM applications. To…

计算机视觉与模式识别 · 计算机科学 2025-10-22 Junjue Wang , Weihao Xuan , Heli Qi , Zhihao Liu , Kunyi Liu , Yuhan Wu , Hongruixuan Chen , Jian Song , Junshi Xia , Zhuo Zheng , Naoto Yokoya

Comprehensive situational awareness is essential for autonomous vehicles operating in safety-critical environments, as it enables the identification and mitigation of potential risks. Although recent Multimodal Large Language Models (MLLMs)…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Sainithin Artham , Shankar Gangisetty , Avijit Dasgupta , C. V. Jawahar

Rapid, fine-grained disaster damage assessment is essential for effective emergency response, yet remains challenging due to limited ground sensors and delays in official reporting. Social media provides a rich, real-time source of…

计算与语言 · 计算机科学 2025-06-05 Zihui Ma , Lingyao Li , Juan Li , Wenyue Hua , Jingxiao Liu , Qingyuan Feng , Yuki Miura

The rapid progress of Multimodal Large Language Models (MLLMs) has unlocked the potential for enhanced 3D scene understanding and spatial reasoning. A recent line of work explores learning spatial reasoning directly from multi-view images,…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Kanghee Lee , Injae Lee , Minseok Kwak , Jungi Hong , Kwonyoung Ryu , Jaesik Park

Natural disasters cause devastating damage to communities and infrastructure every year. Effective disaster response is hampered by the difficulty of accessing affected areas during and after events. Remote sensing has allowed us to monitor…

计算机视觉与模式识别 · 计算机科学 2025-07-23 Shreelekha Revankar , Utkarsh Mall , Cheng Perng Phoo , Kavita Bala , Bharath Hariharan

Timely interpretation of satellite imagery is critical for disaster response, yet existing vision-language benchmarks for remote sensing largely focus on coarse labels and image-level recognition, overlooking the functional understanding…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Sara Tehrani , Yonghao Xu , Leif Haglund , Amanda Berg , Michael Felsberg

The Contrastive Language-Image Pre-training (CLIP) framework has become a widely used approach for multimodal representation learning, particularly in image-text retrieval and clustering. However, its efficacy is constrained by three key…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Tiancheng Gu , Kaicheng Yang , Ziyong Feng , Xingjun Wang , Yanzhao Zhang , Dingkun Long , Yingda Chen , Weidong Cai , Jiankang Deng

Embodied scene understanding requires not only comprehending visual-spatial information that has been observed but also determining where to explore next in the 3D physical world. Existing 3D Vision-Language (3D-VL) models primarily focus…

计算机视觉与模式识别 · 计算机科学 2025-07-31 Ziyu Zhu , Xilin Wang , Yixuan Li , Zhuofan Zhang , Xiaojian Ma , Yixin Chen , Baoxiong Jia , Wei Liang , Qian Yu , Zhidong Deng , Siyuan Huang , Qing Li

Vision-language models enable the understanding and reasoning of complex traffic scenarios through multi-source information fusion, establishing it as a core technology for autonomous driving. However, existing vision-language models are…

计算机视觉与模式识别 · 计算机科学 2025-12-17 Minghui Hou , Wei-Hsing Huang , Shaofeng Liang , Daizong Liu , Tai-Hao Wen , Gang Wang , Runwei Guan , Weiping Ding

The application of generalist multimodal models (GMMs) to specialized scientific domains remains limited due to the scarcity of comprehensive domain-specific datasets that integrate multiple data modalities beyond text and images. In…

In large-scale disaster events, the planning of optimal rescue routes depends on the object detection ability at the disaster scene, with one of the main challenges being the presence of dense and occluded objects. Existing methods, which…

计算机视觉与模式识别 · 计算机科学 2024-05-15 Xin Wu , Zhanchao Huang , Li Wang , Jocelyn Chanussot , Jiaojiao Tian

Multimodal semantic understanding often has to deal with uncertainty, which means the obtained messages tend to refer to multiple targets. Such uncertainty is problematic for our interpretation, including inter- and intra-modal uncertainty.…

计算机视觉与模式识别 · 计算机科学 2023-07-21 Yatai Ji , Junjie Wang , Yuan Gong , Lin Zhang , Yanru Zhu , Hongfa Wang , Jiaxing Zhang , Tetsuya Sakai , Yujiu Yang

The use of robotics in humanitarian demining increasingly involves computer vision techniques to improve landmine detection capabilities. However, in the absence of diverse and realistic datasets, the reliable validation of algorithms…

Visual commonsense understanding requires Vision Language (VL) models to not only understand image and text but also cross-reference in-between to fully integrate and achieve comprehension of the visual scene described. Recently, various…

计算机视觉与模式识别 · 计算机科学 2023-10-24 Zhecan Wang , Haoxuan You , Yicheng He , Wenhao Li , Kai-Wei Chang , Shih-Fu Chang

The recent development in multimodal learning has greatly advanced the research in 3D scene understanding in various real-world tasks such as embodied AI. However, most existing studies are facing two common challenges: 1) they are short of…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Xueying Jiang , Lewei Lu , Ling Shao , Shijian Lu

Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in general visual understanding. However, their application to safety-critical driving scenarios remains limited by an inability to…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Tomaso Trinci , Henrique Piñeiro Monteagudo , Leonardo Taccari

Commonsense reasoning in multimodal contexts remains a foundational challenge in artificial intelligence. We introduce Multimodal UNcommonsense(MUN), a benchmark designed to evaluate models' ability to handle scenarios that deviate from…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Yejin Son , Saejin Kim , Dongjun Min , Younjae Yu

Multimodal large language models (MLLMs) are increasingly deployed in open-ended, real-world environments where inputs are messy, underspecified, and not always trustworthy. Unlike curated benchmarks, these settings frequently involve…

人工智能 · 计算机科学 2025-08-26 Qianqi Yan , Hongquan Li , Shan Jiang , Yang Zhao , Xinze Guan , Ching-Chen Kuo , Xin Eric Wang
‹ 上一页 1 2 3 10 下一页 ›