中文
相关论文

相关论文: From Pixels to Semantics: A Multi-Stage AI Framewo…

200 篇论文

Semantic retrieval of remote sensing (RS) images is a critical task fundamentally challenged by the \textquote{semantic gap}, the discrepancy between a model's low-level visual features and high-level human concepts. While large…

计算机视觉与模式识别 · 计算机科学 2025-12-12 J. Xiao , Y. Guo , X. Zi , K. Thiyagarajan , C. Moreira , M. Prasad

Pre-trained Vision-Language Models (VLMs) utilizing extensive image-text paired data have demonstrated unprecedented image-text association capabilities, achieving remarkable results across various downstream tasks. A critical challenge is…

计算机视觉与模式识别 · 计算机科学 2024-09-10 Zilun Zhang , Tiancheng Zhao , Yulong Guo , Jianwei Yin

Ship detection in remote sensing imagery is a critical task with wide-ranging applications, such as maritime activity monitoring, shipping logistics, and environmental studies. However, existing methods often struggle to capture…

计算机视觉与模式识别 · 计算机科学 2025-08-25 Jiahao Li , Jiancheng Pan , Yuze Sun , Xiaomeng Huang

The increasing frequency and intensity of natural disasters call for rapid and accurate damage assessment. In response, disaster benchmark datasets from high-resolution satellite imagery have been constructed to develop methods for…

计算机视觉与模式识别 · 计算机科学 2025-01-22 Kyeongjin Ahn , Sungwon Han , Sungwon Park , Jihee Kim , Sangyoon Park , Meeyoung Cha

Living in a changing climate, human society now faces more frequent and severe natural disasters than ever before. As a consequence, rapid disaster response during the "Golden 72 Hours" of search and rescue becomes a vital humanitarian…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Hao Li , Liwei Zou , Wenping Yin , Gulsen Taskin , Naoto Yokoya , Danfeng Hong , Wufan Zhao

Post-disaster inspections are critical to emergency management after earthquakes. The availability of data on the condition of civil infrastructure immediately after an earthquake is of great importance for emergency management.…

信号处理 · 电气工程与系统科学 2020-09-25 Xiao Liang , Seyed Omid Sajedi

Recent advancements in computer vision and deep learning techniques have facilitated notable progress in scene understanding, thereby assisting rescue teams in achieving precise damage assessment. In this paper, we present RescueNet, a…

计算机视觉与模式识别 · 计算机科学 2024-05-20 Maryam Rahnemoonfar , Tashnim Chowdhury , Robin Murphy

Medical image segmentation is more clinically valuable when it supports diagnosis rather than merely producing lesion masks. However, diagnostically relevant lesion cues are often subtle and localized, while existing models may be…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Fengyi Zhang , Xujie Zeng , Mohan Liu , Zengyi Wang , Yalong Jiang

Recent advances have demonstrated that Language Vision Models (LVMs) surpass the existing State-of-the-Art (SOTA) in two-dimensional (2D) computer vision tasks, motivating attempts to apply LVMs to three-dimensional (3D) data. While LVMs…

计算机视觉与模式识别 · 计算机科学 2025-09-29 June Moh Goo , Zichao Zeng , Jan Boehm

Current Visual Simultaneous Localization and Mapping (VSLAM) systems often struggle to create maps that are both semantically rich and easily interpretable. While incorporating semantic scene knowledge aids in building richer maps with…

Critical infrastructure, such as transport networks and bridges, are systematically targeted during wars and suffer damage during extensive natural disasters because it is vital for enabling connectivity and transportation of people and…

To respond to disasters such as earthquakes, wildfires, and armed conflicts, humanitarian organizations require accurate and timely data in the form of damage assessments, which indicate what buildings and population centers have been most…

计算机视觉与模式识别 · 计算机科学 2020-12-01 Jihyeon Lee , Joseph Z. Xu , Kihyuk Sohn , Wenhan Lu , David Berthelot , Izzeddin Gur , Pranav Khaitan , Ke-Wei , Huang , Kyriacos Koupparis , Bernhard Kowatsch

Semantic segmentation (SS) of RSIs enables the fine-grained interpretation of surface features, making it a critical task in RS analysis. With the increasing diversity and volume of RSIs collected by sensors on various platforms,…

计算机视觉与模式识别 · 计算机科学 2026-02-12 Quanwei Liu , Tao Huang , Jiaqi Yang , Wei Xiang

Image explanation has been one of the key research interests in the Deep Learning field. Throughout the years, several approaches have been adopted to explain an input image fed by the user. From detecting an object in a given image to…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Debjyoti Das Adhikary , Aritra Hazra , Partha Pratim Chakrabarti

Rapid and accurate situational awareness is essential for effective response during natural disasters, where delays in analysis can significantly hinder decision-making. Training task-specific models for post-disaster assessment is often…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Armin Zarbaft , Ehsan Karimi , Nhut Le , Maryam Rahnemoonfar

We introduce Image2Struct, a benchmark to evaluate vision-language models (VLMs) on extracting structure from images. Our benchmark 1) captures real-world use cases, 2) is fully automatic and does not require human judgment, and 3) is based…

计算机视觉与模式识别 · 计算机科学 2024-10-31 Josselin Somerville Roberts , Tony Lee , Chi Heem Wong , Michihiro Yasunaga , Yifan Mai , Percy Liang

Vision language models (VLMs) can flexibly address various vision tasks through text interactions. Although successful in semantic understanding, state-of-the-art VLMs including GPT-5 still struggle in understanding 3D from 2D inputs. On…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Zhipeng Cai , Ching-Feng Yeh , Hu Xu , Zhuang Liu , Gregory Meyer , Xinjie Lei , Changsheng Zhao , Shang-Wen Li , Vikas Chandra , Yangyang Shi

Traditional natural disaster response involves significant coordinated teamwork where speed and efficiency are key. Nonetheless, human limitations can delay critical actions and inadvertently increase human and economic losses. Agentic…

多智能体系统 · 计算机科学 2024-11-05 Zhaohui Chen , Elyas Asadi Shamsabadi , Sheng Jiang , Luming Shen , Daniel Dias-da-Costa

This thesis explores a multimodal AI framework for enhancing construction safety through the combined analysis of textual and visual data. In safety-critical environments such as construction sites, accident data often exists in multiple…

人工智能 · 计算机科学 2025-11-21 Islem Sahraoui

Vision language models (VLMs) have shown significant promise in visual grounding for images as well as videos. In medical imaging research, VLMs represent a bridge between object detection and segmentation, and report understanding and…

计算机视觉与模式识别 · 计算机科学 2026-03-23 Andrew Seohwan Yu , Mohsen Hariri , Kunio Nakamura , Mingrui Yang , Xiaojuan Li , Vipin Chaudhary