English
Related papers

Related papers: T2Vs Meet VLMs: A Scalable Multimodal Dataset for …

200 papers

Video descriptions are crucial for blind and low vision (BLV) users to access visual content. However, current artificial intelligence models for generating descriptions often fall short due to limitations in the quality of human…

Computer Vision and Pattern Recognition · Computer Science 2025-03-03 Chaoyu Li , Sid Padmanabhuni , Maryam Cheema , Hasti Seifi , Pooyan Fazli

Video recognition has been advanced in recent years by benchmarks with rich annotations. However, research is still mainly limited to human action or sports recognition - focusing on a highly specific video understanding task and thus…

Computer Vision and Pattern Recognition · Computer Science 2020-12-16 Ali Diba , Mohsen Fayyaz , Vivek Sharma , Manohar Paluri , Jurgen Gall , Rainer Stiefelhagen , Luc Van Gool

In an era of rapidly evolving internet technology, the surge in multimodal content, including videos, has expanded the horizons of online communication. However, the detection of toxic content in this diverse landscape, particularly in…

Artificial Intelligence · Computer Science 2024-07-16 Krishanu Maity , A. S. Poornash , Sriparna Saha , Pushpak Bhattacharyya

The spread of election misinformation and harmful political content conveys misleading narratives and poses a serious threat to democratic integrity. Detecting harmful content at early stages is essential for understanding and potentially…

Human-Computer Interaction · Computer Science 2026-02-24 Qile Wang , Prerana Khatiwada , Carolina Coimbra Vieira , Benjamin E. Bagozzi , Kenneth E. Barner , Matthew Louis Mauriello

Short video platforms, such as YouTube, Instagram, or TikTok, are used by billions of users. These platforms expose users to harmful content, ranging from clickbait or physical harms to hate or misinformation. Yet, we lack a comprehensive…

Computer Vision and Pattern Recognition · Computer Science 2025-04-24 Wonjeong Jo , Magdalena Wojcieszak

Given the central role of charts in scientific, business, and communication contexts, enhancing the chart understanding capabilities of vision-language models (VLMs) has become increasingly critical. A key limitation of existing VLMs lies…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Jiangning Zhu , Yuxing Zhou , Zheng Wang , Juntao Yao , Yima Gu , Yuhui Yuan , Shixia Liu

Vision Large Language Models (VLLMs) represent a significant advancement in artificial intelligence by integrating image-processing capabilities with textual understanding, thereby enhancing user interactions and expanding application…

Computation and Language · Computer Science 2025-05-09 Madhur Jindal , Saurabh Deshpande

Training of autonomous driving systems requires extensive datasets with precise annotations to attain robust performance. Human annotations suffer from imperfections, and multiple iterations are often needed to produce high-quality…

Computer Vision and Pattern Recognition · Computer Science 2026-05-01 Santosh Vasa , Aditi Ramadwar , Jnana Rama Krishna Darabattula , Md Zafar Anwar , Stanislaw Antol , Andrei Vatavu , Thomas Monninger , Sihao Ding

Visual impairment affects hundreds of millions of people worldwide, severely limiting their ability to navigate urban environments safely and independently. While wearable assistive devices offer a promising platform for real-time hazard…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Antoni Valls , Jordi Sanchez-Riera

The virtual content in augmented reality (AR) can introduce misleading or harmful information, leading to semantic misunderstandings or user errors. In this work, we focus on visual information manipulation (VIM) attacks in AR, where…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Yanming Xiu , Maria Gorlatova

Despite significant breakthroughs in video analysis driven by the rapid development of large multimodal models (LMMs), there remains a lack of a versatile evaluation benchmark to comprehensively assess these models' performance in video…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Yunxin Li , Xinyu Chen , Baotian Hu , Longyue Wang , Haoyuan Shi , Min Zhang

Static benchmarks for harmful content detection face limitations in scalability and diversity, and may also be affected by contamination from web-scale pre-training corpora. To address these issues, we propose a framework for synthesizing…

Computation and Language · Computer Science 2026-04-21 Huije Lee , Jisu Shin , Hoyun Song , Changgeon Ko , Jong C. Park

Detecting illicit visual content demands more than image-level NSFW flags; moderators must also know what objects make an image illegal and where those objects occur. We introduce a zero-shot pipeline that simultaneously (i) detects if an…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Sheng Hang , Chaoxiang He , Hongsheng Hu , Hanqing Hu , Bin Benjamin Zhu , Shi-Feng Sun , Dawu Gu , Shuo Wang

Hateful meme detection presents a significant challenge as a multimodal task due to the complexity of interpreting implicit hate messages and contextual cues within memes. Previous approaches have fine-tuned pre-trained vision-language…

Computation and Language · Computer Science 2025-02-18 Ming Shan Hee , Roy Ka-Wei Lee

With the rapid advancement of video generation models such as Sora, video quality assessment (VQA) is becoming increasingly crucial for selecting high-quality videos from large-scale datasets used in pre-training. Traditional VQA methods,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Yanyun Pu , Kehan Li , Zeyi Huang , Zhijie Zhong , Kaixiang Yang

Despite the remarkable performance of foundation vision-language models, the shared representation space for text and vision can also encode harmful label associations detrimental to fairness. While prior work has uncovered bias in…

Computer Vision and Pattern Recognition · Computer Science 2024-04-17 Caner Hazirbas , Alicia Sun , Yonathan Efroni , Mark Ibrahim

Hateful memes are widespread in social media and convey negative information. The main challenge of hateful memes detection is that the expressive meaning can not be well recognized by a single modality. In order to further integrate modal…

Computer Vision and Pattern Recognition · Computer Science 2020-12-10 Weibo Zhang , Guihua Liu , Zhuohua Li , Fuqing Zhu

The automatic identification of harmful content online is of major concern for social media platforms, policymakers, and society. Researchers have studied textual, visual, and audio content, but typically in isolation. Yet, harmful content…

Safety evaluation of multimodal foundation models often treats vision and language inputs separately, missing risks from joint interpretation where benign content becomes harmful in combination. Existing approaches also fail to distinguish…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Shruti Palaskar , Leon Gatys , Mona Abdelrahman , Mar Jacobo , Larry Lindsey , Rutika Moharir , Gunnar Lund , Yang Xu , Navid Shiee , Jeffrey Bigham , Charles Maalouf , Joseph Yitan Cheng

Detecting anomalous hazards in visual data, particularly in video streams, is a critical challenge in autonomous driving. Existing models often struggle with unpredictable, out-of-label hazards due to their reliance on predefined object…

Computer Vision and Pattern Recognition · Computer Science 2025-04-21 Shashank Shriram , Srinivasa Perisetla , Aryan Keskar , Harsha Krishnaswamy , Tonko Emil Westerhof Bossen , Andreas Møgelmose , Ross Greer