English
Related papers

Related papers: ViDAS: Vision-based Danger Assessment and Scoring

200 papers

With the burgeoning growth of online video platforms and the escalating volume of video content, the demand for proficient video understanding tools has intensified markedly. Given the remarkable capabilities of large language models (LLMs)…

Construction safety inspections typically involve a human inspector identifying safety concerns on-site. With the rise of powerful Vision Language Models (VLMs), researchers are exploring their use for tasks such as detecting safety rule…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Xuezheng Chen , Zhengbo Zou

Multi-modal Large language models (MLLMs) show remarkable ability in video understanding. Nevertheless, understanding long videos remains challenging as the models can only process a finite number of frames in a single inference,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Yucheng Suo , Fan Ma , Linchao Zhu , Tianyi Wang , Fengyun Rao , Yi Yang

Automatic detection of natural disasters and incidents has become more important as a tool for fast response. There have been many studies to detect incidents using still images and text. However, the number of approaches that exploit…

Computer Vision and Pattern Recognition · Computer Science 2023-01-10 Duygu Sesver , Alp Eren Gençoğlu , Çağrı Emre Yıldız , Zehra Günindi , Faeze Habibi , Ziya Ata Yazıcı , Hazım Kemal Ekenel

Vision-language models (VLMs) are increasingly applied to identify unsafe or inappropriate images due to their internal ethical standards and powerful reasoning abilities. However, it is still unclear whether they can recognize various…

Cryptography and Security · Computer Science 2025-07-16 Yiting Qu , Michael Backes , Yang Zhang

Multimodal large language models (MLLMs) can simultaneously process visual, textual, and auditory data, capturing insights that complement human analysis. However, existing video question-answering (VidQA) benchmarks and datasets often…

Machine Learning · Computer Science 2024-12-23 Jean Park , Kuk Jin Jang , Basam Alasaly , Sriharsha Mopidevi , Andrew Zolensky , Eric Eaton , Insup Lee , Kevin Johnson

The last two years have seen a rapid growth in concerns around the safety of large language models (LLMs). Researchers and practitioners have met these concerns by creating an abundance of datasets for evaluating and improving LLM safety.…

Computation and Language · Computer Science 2025-01-13 Paul Röttger , Fabio Pernisi , Bertie Vidgen , Dirk Hovy

Advertisement videos serve as a rich and valuable source of purpose-driven information, encompassing high-quality visual, textual, and contextual cues designed to engage viewers. They are often more complex than general videos of similar…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Zheyuan Zhang , Monica Dou , Linkai Peng , Hongyi Pan , Ulas Bagci , Boqing Gong

Ensuring the safety of vulnerable road users (VRUs), such as pedestrians and cyclists, is a critical challenge for autonomous driving systems, as crashes involving VRUs often result in severe or fatal consequences. While multimodal large…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Younggun Kim , Ahmed S. Abdelrahman , Mohamed Abdel-Aty

The advent of large vision-language models (LVLMs) has spurred research into their applications in multi-modal contexts, particularly in video understanding. Traditional VideoQA benchmarks, despite providing quantitative metrics, often fail…

Computer Vision and Pattern Recognition · Computer Science 2024-10-31 Xinyu Fang , Kangrui Mao , Haodong Duan , Xiangyu Zhao , Yining Li , Dahua Lin , Kai Chen

Large Language Models integrating textual and visual inputs have introduced new possibilities for interpreting complex data. Despite their remarkable ability to generate coherent and contextually relevant text based on visual stimuli, the…

Human-Computer Interaction · Computer Science 2025-01-08 Giulio Antonio Abbo , Tony Belpaeme

We introduce RoadSocial, a large-scale, diverse VideoQA dataset tailored for generic road event understanding from social media narratives. Unlike existing datasets limited by regional bias, viewpoint bias and expert-driven annotations,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Chirag Parikh , Deepti Rawat , Rakshitha R. T. , Tathagata Ghosh , Ravi Kiran Sarvadevabhatla

Understanding surveillance video content remains a critical yet underexplored challenge in vision-language research, particularly due to its real-world complexity, irregular event dynamics, and safety-critical implications. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Bo Liu , Pengfei Qiao , Minhan Ma , Xuange Zhang , Yinan Tang , Peng Xu , Kun Liu , Tongtong Yuan

As a leading online platform with a vast global audience, YouTube's extensive reach also makes it susceptible to hosting harmful content, including disinformation and conspiracy theories. This study explores the use of open-weight Large…

Computation and Language · Computer Science 2025-07-08 Leonardo La Rocca , Francesco Corso , Francesco Pierri

Large language models (LLMs) are increasingly applied in mental health support systems, where reliable recognition of high-risk states such as suicidal ideation and self-harm is safety-critical. However, existing evaluations primarily rely…

Artificial Intelligence · Computer Science 2026-03-12 Yihe Zhang , Cheyenne N Mohawk , Kaiying Han , Vijay Srinivas Tida , Manyu Li , Xiali Hei

The proliferation of Vision-Language Models (VLMs) introduces profound privacy risks from personal videos. This paper addresses the critical yet unexplored inferential privacy threat, the risk of inferring sensitive personal attributes over…

Human-Computer Interaction · Computer Science 2025-11-05 Shuning Zhang , Zhaoxin Li , Changxi Wen , Ying Ma , Simin Li , Gengrui Zhang , Ziyi Zhang , Yibo Meng , Hantao Zhao , Xin Yi , Hewu Li

Towards open-ended Video Anomaly Detection (VAD), existing methods often exhibit biased detection when faced with challenging or unseen events and lack interpretability. To address these drawbacks, we propose Holmes-VAD, a novel framework…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Huaxin Zhang , Xiaohao Xu , Xiang Wang , Jialong Zuo , Chuchu Han , Xiaonan Huang , Changxin Gao , Yuehuan Wang , Nong Sang

With the rapid evolution of large language models (LLMs), there is a growing concern that they may pose risks or have negative social impacts. Therefore, evaluation of human values alignment is becoming increasingly important. Previous work…

Computation and Language · Computer Science 2023-07-20 Guohai Xu , Jiayi Liu , Ming Yan , Haotian Xu , Jinghui Si , Zhuoran Zhou , Peng Yi , Xing Gao , Jitao Sang , Rong Zhang , Ji Zhang , Chao Peng , Fei Huang , Jingren Zhou

Benefiting from the powerful capabilities of Large Language Models (LLMs), pre-trained visual encoder models connected to an LLMs can realize Vision Language Models (VLMs). However, existing research shows that the visual modality of VLMs…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Zhendong Liu , Yuanbi Nie , Yingshui Tan , Xiangyu Yue , Qiushi Cui , Chongjun Wang , Xiaoyong Zhu , Bo Zheng

Recently, Large Vision-Language Models (LVLMs) have made significant strides across diverse multimodal tasks and benchmarks. This paper reveals a largely under-explored problem from existing video-involved LVLMs - language bias, where…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Yiming Yang , Yangyang Guo , Hui Lu , Yan Wang