English
Related papers

Related papers: Contextual Cross-Modal Attention for Audio-Visual …

200 papers

Deepfakes are realistic face manipulations that can pose serious threats to security, privacy, and trust. Existing methods mostly treat this task as binary classification, which uses digital labels or mask signals to train the detection…

Computer Vision and Pattern Recognition · Computer Science 2024-02-08 Ke Sun , Shen Chen , Taiping Yao , Haozhe Yang , Xiaoshuai Sun , Shouhong Ding , Rongrong Ji

Deepfakes have raised significant concerns due to their potential to spread false information and compromise digital media integrity. Current deepfake detection models often struggle to generalize across a diverse range of deepfake…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Deressa Wodajo Deressa , Hannes Mareen , Peter Lambert , Solomon Atnafu , Zahid Akhtar , Glenn Van Wallendael

In this paper we propose a novel human-centered approach for detecting forgery in face images, using dynamic prototypes as a form of visual explanations. Currently, most state-of-the-art deepfake detections are based on black-box models…

Computer Vision and Pattern Recognition · Computer Science 2021-01-18 Loc Trinh , Michael Tsang , Sirisha Rambhatla , Yan Liu

Audio-text retrieval enables semantic alignment between audio content and natural language queries, supporting applications in multimedia search, accessibility, and surveillance. However, current state-of-the-art approaches struggle with…

Computation and Language · Computer Science 2026-04-28 Meizhu Liu , Matthew Rowe , Amit Agarwal , Michael Avendi , Yassi Abbasi , Hitesh Laxmichand Patel , Paul Li , Kyu J. Han , Tao Sheng , Sujith Ravi , Dan Roth

Existing deepfake detection methods often exhibit bias, lack transparency, and fail to capture temporal information, leading to biased decisions and unreliable results across different demographic groups. In this paper, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Akihito Yoshii , Ryosuke Sonoda , Ramya Srinivasan

Deep neural network approaches have demonstrated high performance in object recognition (CNN) and detection (Faster-RCNN) tasks, but experiments have shown that such architectures are vulnerable to adversarial attacks (FFF, UAP): low…

Computer Vision and Pattern Recognition · Computer Science 2020-11-16 Faisal Alamri , Sinan Kalkan , Nicolas Pugeault

Audio-visual target speech extraction, which aims to extract a certain speaker's speech from the noisy mixture by looking at lip movements, has made significant progress combining time-domain speech separation models and visual feature…

Multimedia · Computer Science 2023-03-07 Zhongweiyang Xu , Xulin Fan , Mark Hasegawa-Johnson

Deepfakes are the result of digital manipulation to forge realistic yet fake imagery. With the astonishing advances in deep generative models, fake images or videos are nowadays obtained using variational autoencoders (VAEs) or Generative…

Computer Vision and Pattern Recognition · Computer Science 2022-06-29 Davide Coccomini , Nicola Messina , Claudio Gennaro , Fabrizio Falchi

Cross-modality fusing complementary information of multispectral remote sensing image pairs can improve the perception ability of detection algorithms, making them more robust and reliable for a wider range of applications, such as…

Computer Vision and Pattern Recognition · Computer Science 2021-12-07 Qingyun Fang , Zhaokui Wang

Deepfake detection refers to detecting artificially generated or edited faces in images or videos, which plays an essential role in visual information security. Despite promising progress in recent years, Deepfake detection remains a…

Computer Vision and Pattern Recognition · Computer Science 2023-12-11 Chunlei Peng , Huiqing Guo , Decheng Liu , Nannan Wang , Ruimin Hu , Xinbo Gao

Multimedia data, particularly images and videos, is integral to various applications, including surveillance, visual interaction, biometrics, evidence gathering, and advertising. However, amateur or skilled counterfeiters can simulate them…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Kutub Uddin , Nusrat Tasnim , Byung Tae Oh

Pioneering advancements in artificial intelligence, especially in genAI, have enabled significant possibilities for content creation, but also led to widespread misinformation and false content. The growing sophistication and realism of…

Artificial Intelligence · Computer Science 2024-11-14 Dinesh Srivasthav P , Badri Narayan Subudhi

Audio tagging aims to perform multi-label classification on audio chunks and it is a newly proposed task in the Detection and Classification of Acoustic Scenes and Events 2016 (DCASE 2016) challenge. This task encourages research efforts to…

Sound · Computer Science 2017-03-20 Yong Xu , Qiuqiang Kong , Qiang Huang , Wenwu Wang , Mark D. Plumbley

The spread of misinformation through synthetically generated yet realistic images and videos has become a significant problem, calling for robust manipulation detection methods. Despite the predominant effort of detecting face manipulation…

Computer Vision and Pattern Recognition · Computer Science 2019-05-17 Ekraam Sabir , Jiaxin Cheng , Ayush Jaiswal , Wael AbdAlmageed , Iacopo Masi , Prem Natarajan

Multimodal generative models are rapidly evolving, leading to a surge in the generation of realistic video and audio that offers exciting possibilities but also serious risks. Deepfake videos, which can convincingly impersonate individuals,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-22 Hannah Lee , Changyeon Lee , Kevin Farhat , Lin Qiu , Steve Geluso , Aerin Kim , Oren Etzioni

The rapid advancement of deepfake generation techniques poses significant threats to public safety and causes societal harm through the creation of highly realistic synthetic facial media. While existing detection methods demonstrate…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Jianfeng Liao , Yichen Wei , Raymond Chan Ching Bon , Shulan Wang , Kam-Pui Chow , Kwok-Yan Lam

Audio-visual quality assessment (AVQA) is essential for streaming, teleconferencing, and immersive media. In realistic streaming scenarios, distortions are often asymmetric, where one modality may be severely degraded while the other…

Multimedia · Computer Science 2026-05-05 Mayesha Maliha R. Mithila , Mylene C. Q. Farias

In the contemporary digital age, the proliferation of deepfakes presents a formidable challenge to the sanctity of information dissemination. Audio deepfakes, in particular, can be deceptively realistic, posing significant risks in…

Sound · Computer Science 2024-02-28 Karthik Sivarama Krishnan , Koushik Sivarama Krishnan

This report presents our approach for the IEEE SP Cup 2025: Deepfake Face Detection in the Wild (DFWild-Cup), focusing on detecting deepfakes across diverse datasets. Our methodology employs advanced backbone models, including MaxViT,…

Over the last years, there has been an unprecedented proliferation of fake news. As a consequence, we are more susceptible to the pernicious impact that misinformation and disinformation spreading can have in different segments of our…

Computation and Language · Computer Science 2021-12-10 Santiago Alonso-Bartolome , Isabel Segura-Bedmar