English
Related papers

Related papers: A Unified Framework for Modality-Agnostic Deepfake…

200 papers

Generative AI advances rapidly, allowing the creation of very realistic manipulated video and audio. This progress presents a significant security and ethical threat, as malicious users can exploit DeepFake techniques to spread…

Multimedia · Computer Science 2025-06-09 Marcel Klemt , Carlotta Segna , Anna Rohrbach

Active speaker detection requires a solid integration of multi-modal cues. While individual modalities can approximate a solution, accurate predictions can only be achieved by explicitly fusing the audio and visual features and modeling…

Computer Vision and Pattern Recognition · Computer Science 2021-10-06 Juan León-Alcázar , Fabian Caba Heilbron , Ali Thabet , Bernard Ghanem

In today's era of digital misinformation, we are increasingly faced with new threats posed by video falsification techniques. Such falsifications range from cheapfakes (e.g., lookalikes or audio dubbing) to deepfakes (e.g., sophisticated AI…

Computer Vision and Pattern Recognition · Computer Science 2022-12-05 Shruti Agarwal , Liwen Hu , Evonne Ng , Trevor Darrell , Hao Li , Anna Rohrbach

Audio DeepFakes allow the creation of high-quality, convincing utterances and therefore pose a threat due to its potential applications such as impersonation or fake news. Methods for detecting these manipulations should be characterized by…

Sound · Computer Science 2022-10-13 Piotr Kawa , Marcin Plata , Piotr Syga

Multimodal deepfakes are proliferating on social media and threaten authenticity, information integrity, and digital forensics. Existing benchmarks are constrained by their single-modality scope, simplified manipulations, or unrealistic…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Tianxiao Li , Zhenglin Huang , Haiquan Wen , Yiwei He , Xinze Li , Bingyu Zhu , Wuhui Duan , Congang Chen , Zeyu Fu , Yi Dong , Baoyuan Wu , Jason Li , Guangliang Cheng

The advancements of AI-synthesized human voices have introduced a growing threat of impersonation and disinformation. It is therefore of practical importance to developdetection methods for synthetic human voices. This work proposes a new…

Sound · Computer Science 2023-04-28 Chengzhe Sun , Shan Jia , Shuwei Hou , Ehab AlBadawy , Siwei Lyu

The rapid advancement of Deepfake technologies and video manipulation tools poses a critical challenge to multimedia forensics, judicial evidence integrity, and information authenticity. Current detectors rely on single-modality signals,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Hoda Osama Elkhodary , Sherin Mostafa Youssef , Marwa Elshenawy , Dalia Sobhy

The rapid advancement of AI-generated multimodal video-audio content has raised significant concerns regarding information security and content authenticity. Existing synthetic video datasets predominantly focus on the visual modality…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Mengxue Hu , Yunfeng Diao , Changtao Miao , Zhiqing Guo , Jianshu Li , Zhe Li , Joey Tianyi Zhou

It is now well established from a variety of studies that there is a significant benefit from combining video and audio data in detecting active speakers. However, either of the modalities can potentially mislead audiovisual fusion by…

Audio-visual segmentation (AVS) aims to segment objects in videos based on audio cues. Existing AVS methods are primarily designed to enhance interaction efficiency but pay limited attention to modality representation discrepancies and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Mingfeng Zha , Tianyu Li , Guoqing Wang , Peng Wang , Yangyang Wu , Yang Yang , Heng Tao Shen

Deepfakes generated by advanced generative models have rapidly posed serious threats, yet existing audiovisual deepfake detection approaches struggle to generalize to unseen manipulation methods. To address this, we propose a novel…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Hyemin Boo , Eunsang Lee , Jiyoung Lee

Automated deception detection is crucial for assisting humans in accurately assessing truthfulness and identifying deceptive behavior. Conventional contact-based techniques, like polygraph devices, rely on physiological signals to determine…

Previous deepfake detection methods mostly depend on low-level textural features vulnerable to perturbations and fall short of detecting unseen forgery methods. In contrast, high-level semantic features are less susceptible to perturbations…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Ziyuan Fang , Hanqing Zhao , Tianyi Wei , Wenbo Zhou , Ming Wan , Zhanyi Wang , Weiming Zhang , Nenghai Yu

The threat of Audio-Video (AV) forgery is rapidly evolving beyond human-centric deepfakes to include more diverse manipulations across complex natural scenes. However, existing benchmarks are still confined to DeepFake-based forgeries and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Shuhan Xia , Peipei Li , Xuannan Liu , Dongsen Zhang , Xinyu Guo , Zekun Li

Automatic deception detection is an important task that has gained momentum in computational linguistics due to its potential applications. In this paper, we propose a simple yet tough to beat multi-modal neural model for deception…

Computation and Language · Computer Science 2018-03-21 Gangeshwar Krishnamurthy , Navonil Majumder , Soujanya Poria , Erik Cambria

With recent advances in speech synthesis including text-to-speech (TTS) and voice conversion (VC) systems enabling the generation of ultra-realistic audio deepfakes, there is growing concern about their potential misuse. However, most…

Sound · Computer Science 2024-04-24 Zuheng Kang , Yayun He , Botao Zhao , Xiaoyang Qu , Junqing Peng , Jing Xiao , Jianzong Wang

Self-supervised representations excel at many vision and speech tasks, but their potential for audio-visual deepfake detection remains underexplored. Unlike prior work that uses these features in isolation or buried within complex…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Dragos-Alexandru Boldisor , Stefan Smeu , Dan Oneata , Elisabeta Oneata

Audio deepfake detection has become increasingly challenging due to rapid advances in speech synthesis and voice conversion technologies, particularly under channel distortions, replay attacks, and real-world recording conditions. This…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-13 K. A. Shahriar

Deepfake technology has raised concerns about the authenticity of digital content, necessitating the development of effective detection methods. However, the widespread availability of deepfakes has given rise to a new challenge in the form…

Computer Vision and Pattern Recognition · Computer Science 2024-03-15 Sarwar Khan

Generative models achieve remarkable results in multiple data domains, including images and texts, among other examples. Unfortunately, malicious users exploit synthetic media for spreading misinformation and disseminating deepfakes.…

Artificial Intelligence · Computer Science 2025-08-04 Tom Or , Omri Azencot