基于大语言模型和视觉语言模型的建筑工地自动危险源检测
摘要
本论文探讨了一种多模态AI框架,用于通过结合文本和视觉数据分析来增强建筑工地的安全性。在安全关键环境中,如建筑工地,事故数据常以多种格式存在,如书面报告、检查记录和现场图像,这使得使用传统方法综合危险源具有挑战性。为此,本论文提出了一种结合文本和图像分析的多模态AI框架,以协助识别建筑工地上的安全危险源。 conducted two case studies to evaluate the capabilities of large language models (LLMs) and vision-language models (VLMs) for automated hazard identification. The first case study introduces a hybrid pipeline that utilizes GPT 4o and GPT 4o mini to extract structured insights from a dataset of 28,000 OSHA accident reports (2000-2025). The second case study extends this investigation using Molmo 7B and Qwen2 VL 2B, lightweight, open-source VLMs. Using the public ConstructionSite10k dataset, the performance of the two models was evaluated on rule-level safety violation detection using natural language prompts. This experiment served as a cost-aware benchmark against proprietary models and allowed testing at scale with ground-truth labels. Despite their smaller size, Molmo 7B and Quen2 VL 2B showed competitive performance in certain prompt configurations, reinforcing the feasibility of low-resource multimodal systems for rule-aware safety monitoring.
引用
@article{arxiv.2511.15720,
title = {Automated Hazard Detection in Construction Sites Using Large Language and Vision-Language Models},
author = {Islem Sahraoui},
journal= {arXiv preprint arXiv:2511.15720},
year = {2025}
}
备注
Master thesis, University of Houton