English
Related papers

Related papers: VidHarm: A Clip Based Dataset for Harmful Content …

200 papers

Hate speech is a pressing issue in modern society, with significant effects both online and offline. Recent research in hate speech detection has primarily centered on text-based media, largely overlooking multimodal content such as videos.…

Multimedia · Computer Science 2024-08-13 Han Wang , Tan Rui Yang , Usman Naseem , Roy Ka-Wei Lee

This paper introduces VideoMind, a video-centric omni-modal dataset designed for deep video content cognition and enhanced multi-modal feature representation. The dataset comprises 103K video samples (3K reserved for testing), each paired…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Baoyao Yang , Wanyun Li , Dixin Chen , Junxiang Chen , Wenbin Yao , Haifeng Lin

Automatic video summarization is still an unsolved problem due to several challenges. The currently available datasets either have very short videos or have few long videos of only a particular type. We introduce a new benchmarking video…

Computer Vision and Pattern Recognition · Computer Science 2021-01-27 Vishal Kaushal , Suraj Kothawade , Anshul Tomar , Rishabh Iyer , Ganesh Ramakrishnan

Recent developments in vision-language models have significantly advanced video understanding. Existing datasets and tasks, however, have notable limitations. Most datasets are confined to short videos with limited events and narrow…

Computer Vision and Pattern Recognition · Computer Science 2025-01-13 Ridouane Ghermi , Xi Wang , Vicky Kalogeiton , Ivan Laptev

Movie highlights stand out of the screenplay for efficient browsing and play a crucial role on social media platforms. Based on existing efforts, this work has two observations: (1) For different annotators, labeling highlight has…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Bei Gan , Xiujun Shu , Ruizhi Qiao , Haoqian Wu , Keyu Chen , Hanjun Li , Bo Ren

Although recent traffic benchmarks have advanced multimodal data analysis, they generally lack systematic evaluation aligned with official safety standards. To fill this gap, we introduce RoadSafe365, a large-scale vision-language benchmark…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Xinyu Liu , Darryl C. Jacob , Yuxin Liu , Xinsong Du , Muchao Ye , Bolei Zhou , Pan He

As information becomes more accessible, user-generated videos are increasing in length, placing a burden on viewers to sift through vast content for valuable insights. This trend underscores the need for an algorithm to extract key video…

Computer Vision and Pattern Recognition · Computer Science 2024-12-13 Lingfeng Yang , Zhenyuan Chen , Xiang Li , Peiyang Jia , Liangqu Long , Jian Yang

High-quality and consistent annotations are fundamental to the successful development of robust machine learning models. Traditional data annotation methods are resource-intensive and inefficient, often leading to a reliance on third-party…

Computer Vision and Pattern Recognition · Computer Science 2024-02-12 Amir Ziai , Aneesh Vartakavi

Detecting transitions between intro/credits and main content in videos is a crucial task for content segmentation, indexing, and recommendation systems. Manual annotation of such transitions is labor-intensive and error-prone, while…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Vasilii Korolkov , Andrey Yanchenko

Vision Language Models (VLMs) are poised to revolutionize the digital transformation of pharmacyceutical industry by enabling intelligent, scalable, and automated multi-modality content processing. Traditional manual annotation of…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Suyash Mishra , Qiang Li , Srikanth Patil , Anubhav Girdhar

To help customers make better-informed viewing choices, video-streaming services try to moderate their content and provide more visibility into which portions of their movies and TV episodes contain age-appropriate material (e.g., nudity,…

Computer Vision and Pattern Recognition · Computer Science 2022-06-20 Xiang Hao , Jingxiang Chen , Shixing Chen , Ahmed Saad , Raffay Hamid

Constructing supervised machine learning models for real-world video analysis require substantial labeled data, which is costly to acquire due to scarce domain expertise and laborious manual inspection. While data programming shows promise…

Computer Vision and Pattern Recognition · Computer Science 2023-11-02 Jianben He , Xingbo Wang , Kam Kwai Wong , Xijie Huang , Changjian Chen , Zixin Chen , Fengjie Wang , Min Zhu , Huamin Qu

Retrieving target videos based on text descriptions is a task of great practical value and has received increasing attention over the past few years. Despite recent progress, imperfect annotations in existing video retrieval datasets have…

Computer Vision and Pattern Recognition · Computer Science 2022-07-22 Zeyu Wang , Yu Wu , Karthik Narasimhan , Olga Russakovsky

Video description involves the generation of the natural language description of actions, events, and objects in the video. There are various applications of video description by filling the gap between languages and vision for visually…

Computer Vision and Pattern Recognition · Computer Science 2020-12-01 Alok Singh , Thoudam Doren Singh , Sivaji Bandyopadhyay

Automatic detection of natural disasters and incidents has become more important as a tool for fast response. There have been many studies to detect incidents using still images and text. However, the number of approaches that exploit…

Computer Vision and Pattern Recognition · Computer Science 2023-01-10 Duygu Sesver , Alp Eren Gençoğlu , Çağrı Emre Yıldız , Zehra Günindi , Faeze Habibi , Ziya Ata Yazıcı , Hazım Kemal Ekenel

We present Audiovisual Moments in Time (AVMIT), a large-scale dataset of audiovisual action events. In an extensive annotation task 11 participants labelled a subset of 3-second audiovisual videos from the Moments in Time dataset (MIT). For…

Machine Learning · Computer Science 2023-08-21 Michael Joannou , Pia Rotshtein , Uta Noppeney

Video description is the automatic generation of natural language sentences that describe the contents of a given video. It has applications in human-robot interaction, helping the visually impaired and video subtitling. The past few years…

Computer Vision and Pattern Recognition · Computer Science 2020-03-04 Nayyer Aafaq , Ajmal Mian , Wei Liu , Syed Zulqarnain Gilani , Mubarak Shah

We introduce ViSMap: Unsupervised Video Summarisation by Meta Prompting, a system to summarise hour long videos with no-supervision. Most existing video understanding models work well on short videos of pre-segmented events, yet they…

Computer Vision and Pattern Recognition · Computer Science 2025-04-23 Jian Hu , Dimitrios Korkinof , Shaogang Gong , Mariano Beguerisse-Diaz

Large vision-language models (LVLMs) are increasingly used for tasks where detecting multimodal harmful content is crucial, such as online content moderation. However, real-world harmful content is often camouflaged, relying on nuanced…

Multimedia · Computer Science 2025-12-04 Yanhui Li , Qi Zhou , Zhihong Xu , Huizhong Guo , Wenhai Wang , Dongxia Wang

Vision-Language Models (VLMs) are able to process increasingly longer videos. Yet, important visual information is easily lost throughout the entire context and missed by VLMs. Also, it is important to design tools that enable…

Computation and Language · Computer Science 2026-01-09 Galann Pennec , Zhengyuan Liu , Nicholas Asher , Philippe Muller , Nancy F. Chen