English
Related papers

Related papers: VidHarm: A Clip Based Dataset for Harmful Content …

200 papers

We introduce the Lecture Video Visual Objects (LVVO) dataset, a new benchmark for visual object detection in educational video content. The dataset consists of 4,000 frames extracted from 245 lecture videos spanning biology, computer…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Dipayan Biswas , Shishir Shah , Jaspal Subhlok

In the recent years, there has been a tremendous increase in the amount of video content uploaded to social networking and video sharing websites like Facebook and Youtube. As of result of this, the risk of children getting exposed to adult…

Computer Vision and Pattern Recognition · Computer Science 2021-09-21 Praveen Tirupattur , Christian Schulze , Andreas Dengel

Recent advances in multimodal AI have enabled progress in detecting synthetic and out-of-context content. However, existing efforts largely overlook the intent behind AI-generated images. To fill this gap, we introduce S-HArM, a multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-09-10 Anastasios Skoularikis , Stefanos-Iordanis Papadopoulos , Symeon Papadopoulos , Panagiotis C. Petrantonakis

Long-form video understanding remains challenging for Vision-Language Models (VLMs) due to the inherent tension between computational constraints and the need to capture information distributed across thousands of frames. Existing…

Computer Vision and Pattern Recognition · Computer Science 2026-02-05 Junbo Zou , Ziheng Huang , Shengjie Zhang , Liwen Zhang , Weining Shen

Recent advances in Generative AI (GenAI) have led to significant improvements in the quality of generated visual content. As AI-generated visual content becomes increasingly indistinguishable from real content, the challenge of detecting…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Keerthi Veeramachaneni , Praveen Tirupattur , Amrit Singh Bedi , Mubarak Shah

This paper introduces a new challenge and datasets to foster research toward designing systems that can understand medical videos and provide visual answers to natural language questions. We believe medical videos may provide the best…

Computer Vision and Pattern Recognition · Computer Science 2022-02-01 Deepak Gupta , Kush Attal , Dina Demner-Fushman

Lectures are a learning experience for both students and teachers. Students learn from teachers about the subject material, while teachers learn from students about how to refine their instruction. However, online student feedback is…

Computation and Language · Computer Science 2023-06-16 Rose E. Wang , Pawan Wirawarn , Noah Goodman , Dorottya Demszky

Video Large Language Models (VideoLLMs) are increasingly deployed on numerous critical applications, where users rely on auto-generated summaries while casually skimming the video stream. We show that this interaction hides a critical…

Multimedia · Computer Science 2025-11-18 Yuxin Cao , Wei Song , Derui Wang , Jingling Xue , Jin Song Dong

The standard way of training video models entails sampling at each iteration a single clip from a video and optimizing the clip prediction with respect to the video-level label. We argue that a single clip may not have enough temporal…

Computer Vision and Pattern Recognition · Computer Science 2021-04-06 Xitong Yang , Haoqi Fan , Lorenzo Torresani , Larry Davis , Heng Wang

In this paper we introduce a new dataset for 360-degree video summarization: the transformation of 360-degree video content to concise 2D-video summaries that can be consumed via traditional devices, such as TV sets and smartphones. The…

Computer Vision and Pattern Recognition · Computer Science 2024-06-06 Ioannis Kontostathis , Evlampios Apostolidis , Vasileios Mezaris

Text-to-video generation has surged in interest since Sora, yet open-source models still face a data bottleneck: there is no large, high-quality, easily obtainable video-text corpus. Existing public datasets typically require manual YouTube…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Timing Yang , Sucheng Ren , Alan Yuille , Feng Wang

Adult content detection still poses a great challenge for automation. Existing classifiers primarily focus on distinguishing between erotic and non-erotic texts. However, they often need more nuance in assessing the potential harm.…

Computation and Language · Computer Science 2023-10-24 Inez Okulska , Emilia Wiśnios

To open up new possibilities to assess the multimodal perceptual quality of omnidirectional media formats, we proposed a novel open source 360 audiovisual (AV) quality dataset. The dataset consists of high-quality 360 video clips in…

Multimedia · Computer Science 2022-05-18 Randy F Fela , Andréas Pastor , Patrick Le Callet , Nick Zacharov , Toinon Vigier , Søren Forchhammer

Humans are arguably one of the most important subjects in video streams, many real-world applications such as video summarization or video editing workflows often require the automatic search and retrieval of a person of interest. Despite…

Computer Vision and Pattern Recognition · Computer Science 2021-06-04 Juan Leon Alcazar , Long Mai , Federico Perazzi , Joon-Young Lee , Pablo Arbelaez , Bernard Ghanem , Fabian Caba Heilbron

Partially relevant video retrieval aims to retrieve untrimmed videos using text queries that describe only partial content. However, the inherent asymmetry between brief queries and rich video content inevitably introduces uncertainty into…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Jun Li , Peifeng Lai , Xuhang Lou , Jinpeng Wang , Yuting Wang , Ke Chen , Yaowei Wang , Shu-Tao Xia

Detecting customized moments and highlights from videos given natural language (NL) user queries is an important but under-studied topic. One of the challenges in pursuing this direction is the lack of annotated data. To address this issue,…

Computer Vision and Pattern Recognition · Computer Science 2021-11-30 Jie Lei , Tamara L. Berg , Mohit Bansal

Recent advances in AI-generated video have shown strong performance on \emph{text-to-video} tasks, particularly for short clips depicting a single scene. However, current models struggle to generate longer videos with coherent scene…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Hanwen Shen , Jiajie Lu , Yupeng Cao , Xiaonan Yang

Accurate speed estimation of road vehicles is important for several reasons. One is speed limit enforcement, which represents a crucial tool in decreasing traffic accidents and fatalities. Compared with other research areas and domains, the…

Machine Learning · Computer Science 2022-12-06 Slobodan Djukanović , Nikola Bulatović , Ivana Čavor

In this work, we introduce a dataset of video annotated with high quality natural language phrases describing the visual content in a given segment of time. Our dataset is based on the Descriptive Video Service (DVS) that is now encoded on…

Computer Vision and Pattern Recognition · Computer Science 2015-03-04 Atousa Torabi , Christopher Pal , Hugo Larochelle , Aaron Courville

Micro-videos have recently gained immense popularity, sparking critical research in micro-video recommendation with significant implications for the entertainment, advertising, and e-commerce industries. However, the lack of large-scale…

Information Retrieval · Computer Science 2023-09-28 Yongxin Ni , Yu Cheng , Xiangyan Liu , Junchen Fu , Youhua Li , Xiangnan He , Yongfeng Zhang , Fajie Yuan
‹ Prev 1 4 5 6 7 8 10 Next ›