中文
相关论文

相关论文: DisasterM3: A Remote Sensing Vision-Language Datas…

200 篇论文

Aerial imagery is critical for large-scale post-disaster damage assessment. Automated interpretation remains challenging due to clutter, visual variability, and strong cross-event domain shift, while supervised approaches still rely on…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Anna Michailidou , Georgios Angelidis , Vasileios Argyriou , Panagiotis Sarigiannidis , Georgios Th. Papadopoulos

Building damage identification shortly after a disaster is crucial for guiding emergency response and recovery efforts. Although optical satellite imagery is commonly used for disaster mapping, its effectiveness is often hampered by cloud…

计算机视觉与模式识别 · 计算机科学 2025-06-30 Luigi Russo , Deodato Tapete , Silvia Liberata Ullo , Paolo Gamba

Large-scale Vision Language Models (LVLMs) exhibit advanced capabilities in tasks that require visual information, including object detection. These capabilities have promising applications in various industrial domains, such as autonomous…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Haruki Sakajo , Hiroshi Takato , Hiroshi Tsutsui , Komei Soda , Hidetaka Kamigaito , Taro Watanabe

Despite their success, current training pipelines for reasoning VLMs focus on a limited range of tasks, such as mathematical and logical reasoning. As a result, these models face difficulties in generalizing their reasoning capabilities to…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Yuheng Zha , Kun Zhou , Yujia Wu , Yushu Wang , Jie Feng , Zhi Xu , Shibo Hao , Zhengzhong Liu , Eric P. Xing , Zhiting Hu

The widespread use of microblogging platforms like X (formerly Twitter) during disasters provides real-time information to governments and response authorities. However, the data from these platforms is often noisy, requiring automated…

计算与语言 · 计算机科学 2024-12-17 Muhammad Imran , Abdul Wahab Ziaullah , Kai Chen , Ferda Ofli

In all types of disasters, from earthquakes to armed conflicts, aid workers need accurate and timely data such as damage to buildings and population displacement to mount an effective response. Remote sensing provides this data at an…

计算机视觉与模式识别 · 计算机科学 2019-10-16 Joseph Z. Xu , Wenhan Lu , Zebo Li , Pranav Khaitan , Valeriya Zaytseva

Multimodal large language models (MLLMs) have achieved impressive performance across various tasks such as image captioning and visual question answer(VQA); however, they often struggle to accurately interpret depth information inherent in…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Hao Yang , Hongbo Zhang , Yanyan Zhao , Bing Qin

Recently, large vision-language models (LVLMs) unleash powerful analysis capabilities for low Earth orbit (LEO) satellite Earth observation images in the data center. However, fast satellite motion, brief satellite-ground station (GS)…

网络与互联网体系结构 · 计算机科学 2025-07-09 Yuxin Zhang , Jiahao Yang , Zhe Chen , Wenjun Zhu , Jin Zhao , Yue Gao

Recent advances in vision-language models (VLMs) have accelerated their application to indoor safety hazards assessment. However, existing benchmarks suffer from three fundamental limitations: (1) heavy reliance on synthetic datasets…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Qiucheng Yu , Ruijie Xu , Mingang Chen , Xuequan Lu , Jianfeng Dong , Chaochao Lu , Xin Tan

Multimodal Large Language Models (MLLMs) show reasoning promise, yet their visual perception is a critical bottleneck. Strikingly, MLLMs can produce correct answers even while misinterpreting crucial visual elements, masking these…

计算机视觉与模式识别 · 计算机科学 2025-12-11 Aditya Kanade , Tanuja Ganu

We aim to detect and identify multiple objects using multiple cameras and computer vision for disaster response drones. The major challenges are taming detection errors, resolving ID switching and fragmentation, adapting to multi-scale…

计算机视觉与模式识别 · 计算机科学 2022-01-06 Chongkeun Paik , Hyunwoo J. Kim

The increasing rate of road accidents worldwide results not only in significant loss of life but also imposes billions financial burdens on societies. Current research in traffic crash frequency modeling and analysis has predominantly…

计算机视觉与模式识别 · 计算机科学 2024-06-18 Zhiwen Fan , Pu Wang , Yang Zhao , Yibo Zhao , Boris Ivanovic , Zhangyang Wang , Marco Pavone , Hao Frank Yang

Landslides are among the most common natural disasters globally, posing significant threats to human society. Deep learning (DL) has proven to be an effective method for rapidly generating landslide inventories in large-scale disaster…

计算机视觉与模式识别 · 计算机科学 2025-02-28 Guanting Liu , Yi Wang , Xi Chen , Baoyu Du , Penglei Li , Yuan Wu , Zhice Fang

In recent years, social media has emerged as a primary channel for users to promptly share feedback and issues during disasters and emergencies, playing a key role in crisis management. While significant progress has been made in collecting…

计算与语言 · 计算机科学 2025-04-18 Loris Belcastro , Cristian Cosentino , Fabrizio Marozzo , Merve Gündüz-Cüre , Sule Öztürk-Birim

Emergency search and rescue (SAR) operations often require rapid and precise target identification in complex environments where traditional manual drone control is inefficient. In order to address these scenarios, a rapid SAR system,…

This paper introduces a novel benchmark dataset designed to evaluate the capabilities of Vision Language Models (VLMs) on tasks that combine visual reasoning with subject-specific background knowledge in the German language. In contrast to…

人工智能 · 计算机科学 2025-06-30 René Peinl , Vincent Tischler

Human-scene vision-language tasks are increasingly prevalent in diverse social applications, yet recent advancements predominantly rely on models specifically tailored to individual tasks. Emerging research indicates that large…

人工智能 · 计算机科学 2024-11-06 Dawei Dai , Xu Long , Li Yutang , Zhang Yuanhui , Shuyin Xia

We propose SiM3D, the first benchmark considering the integration of multiview and multimodal information for comprehensive 3D anomaly detection and segmentation (ADS), where the task is to produce a voxel-based Anomaly Volume. Moreover,…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Alex Costanzino , Pierluigi Zama Ramirez , Luigi Lella , Matteo Ragaglia , Alessandro Oliva , Giuseppe Lisanti , Luigi Di Stefano

Vision-language model (VLM) fine-tuning for application-specific visual grounding based on natural language instructions has become one of the most popular approaches for learning-enabled autonomous systems. However, such fine-tuning relies…

计算机视觉与模式识别 · 计算机科学 2025-02-03 Joshua R. Waite , Md. Zahid Hasan , Qisai Liu , Zhanhong Jiang , Chinmay Hegde , Soumik Sarkar

Using vision-language models (VLMs) as reward models in reinforcement learning holds promise for reducing costs and improving safety. So far, VLM reward models have only been used for goal-oriented tasks, where the agent must reach a…

计算机视觉与模式识别 · 计算机科学 2024-11-22 Evžen Wybitul , Evan Ryan Gunter , Mikhail Seleznyov , David Lindner