English
Related papers

Related papers: MVAD: A Benchmark Dataset for Multimodal AI-Genera…

200 papers

Existing video anomaly detection datasets are inadequate for representing complex anomalies that occur due to the interactions between objects. The absence of complex anomalies in previous video anomaly detection datasets affects research…

Computer Vision and Pattern Recognition · Computer Science 2025-01-17 Furkan Mumcu , Michael J. Jones , Yasin Yilmaz , Anoop Cherian

The rise of social media and short-form video (SFV) has facilitated a breeding ground for misinformation. With the emergence of large language models, significant research has gone into curbing this misinformation problem with automatic…

Machine Learning · Computer Science 2024-10-23 Andrew Kan , Christopher Kan , Zaid Nabulsi

Video Anomaly Detection (VAD) aims to locate events that deviate from normal patterns in videos. Traditional approaches often rely on extensive labeled data and incur high computational costs. Recent tuning-free methods based on Multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Zhaolin Cai , Fan Li , Ziwei Zheng , Haixia Bi , Lijun He

In this paper, we present a toolchain for a comprehensive audio/video analysis by leveraging deep learning based multimodal approach. To this end, different specific tasks of Speech to Text (S2T), Acoustic Scene Classification (ASC),…

Sound · Computer Science 2024-07-04 Lam Pham , Phat Lam , Tin Nguyen , Hieu Tang , Alexander Schindler

The proliferation of sophisticated AI-generated deepfakes poses critical challenges for digital media authentication and societal security. While existing detection methods perform well within specific generative domains, they exhibit…

Computer Vision and Pattern Recognition · Computer Science 2025-05-26 Naseem Khan , Tuan Nguyen , Amine Bermak , Issa Khalil

The task of deepfakes detection is far from being solved by speech or vision researchers. Several publicly available databases of fake synthetic video and speech were built to aid the development of detection methods. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2023-11-30 Pavel Korshunov , Haolin Chen , Philip N. Garner , Sebastien Marcel

Camouflaged Object Detection (COD) aims to identify objects that blend seamlessly into natural scenes. Although RGB-based methods have advanced, their performance remains limited under challenging conditions. Multispectral imagery,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Yang Li , Tingfa Xu , Shuyan Bai , Peifu Liu , Jianan Li

With the rapid advancement of generative AI, synthetic content across images, videos, and audio has become increasingly realistic, amplifying the risk of misinformation. Existing detection approaches predominantly focus on binary…

Machine Learning · Computer Science 2025-07-23 Xu Yang , Qi Zhang , Shuming Jiang , Yaowen Xu , Zhaofan Zou , Hao Sun , Xuelong Li

Weakly supervised video anomaly detection (WS-VAD) is a crucial area in computer vision for developing intelligent surveillance systems. This system uses three feature streams: RGB video, optical flow, and audio signals, where each stream…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Yuta Kaneko , Abu Saleh Musa Miah , Najmul Hassan , Hyoun-Sup Lee , Si-Woong Jang , Jungpil Shin

Recent progress in generative AI technology has made audio deepfakes remarkably more realistic. While current research on anti-spoofing systems primarily focuses on assessing whether a given audio sample is fake or genuine, there has been…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-26 Nicholas Klein , Tianxiang Chen , Hemlata Tak , Ricardo Casal , Elie Khoury

Generating high-quality cartoon animations multimodal control is challenging due to the complexity of non-human characters, stylistically diverse motions and fine-grained emotions. There is a huge domain gap between real-world videos and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Shuolin Xu , Bingyuan Wang , Zeyu Cai , Fangteng Fu , Yue Ma , Tongyi Lee , Hongchuan Yu , Zeyu Wang

Deepfakes have become a growing concern in recent years, prompting researchers to develop benchmark datasets and detection algorithms to tackle the issue. However, existing datasets suffer from significant drawbacks that hamper their…

Computers and Society · Computer Science 2023-09-07 Beomsang Cho , Binh M. Le , Jiwon Kim , Simon Woo , Shahroz Tariq , Alsharif Abuadbba , Kristen Moore

The rapid advancement of video generation models has enabled the creation of highly realistic synthetic media, raising significant societal concerns regarding the spread of misinformation. However, current detection methods suffer from…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Zhengcen Li , Chenyang Jiang , Hang Zhao , Shiyang Zhou , Yunyang Mo , Feng Gao , Fan Yang , Qiben Shan , Shaocong Wu , Jingyong Su

In recent years, researchers combine both audio and video signals to deal with challenges where actions are not well represented or captured by visual cues. However, how to effectively leverage the two modalities is still under development.…

Computer Vision and Pattern Recognition · Computer Science 2024-01-09 Wentao Zhu

Industrial anomaly detection (IAD) has garnered significant attention and experienced rapid development. However, the recent development of IAD approach has encountered certain difficulties due to dataset limitations. On the one hand, most…

Computer Vision and Pattern Recognition · Computer Science 2024-03-20 Chengjie Wang , Wenbing Zhu , Bin-Bin Gao , Zhenye Gan , Jianning Zhang , Zhihao Gu , Shuguang Qian , Mingang Chen , Lizhuang Ma

Audio and video are two most common modalities in the mainstream media platforms, e.g., YouTube. To learn from multimodal videos effectively, in this work, we propose a novel audio-video recognition approach termed audio video Transformer,…

Computer Vision and Pattern Recognition · Computer Science 2024-01-10 Wentao Zhu

Corporate AI-washing-the strategic misrepresentation of AI capabilities via exaggerated or fabricated cross-channel disclosures-has emerged as a systemic threat to capital market information integrity with the widespread adoption of…

Computers and Society · Computer Science 2026-04-14 Zhanjie Wen , Jingqiao Guo

As Artificial Intelligence (AI) technologies continue to evolve, their use in generating realistic, contextually appropriate content has expanded into various domains. Music, an art form and medium for entertainment, deeply rooted into…

Sound · Computer Science 2024-12-11 Yupei Li , Manuel Milling , Lucia Specia , Björn W. Schuller

Learning associations across modalities is critical for robust multimodal reasoning, especially when a modality may be missing during inference. In this paper, we study this problem in the context of audio-conditioned visual synthesis -- a…

Computer Vision and Pattern Recognition · Computer Science 2020-07-24 Anoop Cherian , Moitreya Chatterjee , Narendra Ahuja

Humans can intuitively infer sounds from silent videos, but whether multimodal large language models can perform modal-mismatch reasoning without accessing target modalities remains relatively unexplored. Current…

Multimedia · Computer Science 2025-05-29 Yong Ren , Chenxing Li , Le Xu , Hao Gu , Duzhen Zhang , Yujie Chen , Manjie Xu , Ruibo Fu , Shan Yang , Dong Yu