English
Related papers

Related papers: Statistics-aware Audio-visual Deepfake Detector

200 papers

We present a learning-based method for detecting real and fake deepfake multimedia content. To maximize information for learning, we extract and analyze the similarity between the two audio and visual modalities from within the same video.…

Computer Vision and Pattern Recognition · Computer Science 2020-08-04 Trisha Mittal , Uttaran Bhattacharya , Rohan Chandra , Aniket Bera , Dinesh Manocha

Recent advances in synthetic speech have made audio deepfakes increasingly realistic, posing significant security risks. Existing detection methods that rely on a single modality, either raw waveform embeddings or spectral based features,…

With the proliferation of deepfake audio, there is an urgent need to investigate their attribution. Current source tracing methods can effectively distinguish in-distribution (ID) categories. However, the rapid evolution of deepfake…

Sound · Computer Science 2024-06-11 Yuankun Xie , Ruibo Fu , Zhengqi Wen , Zhiyong Wang , Xiaopeng Wang , Haonnan Cheng , Long Ye , Jianhua Tao

Existing Audio Deepfake Detection (ADD) systems often struggle to generalise effectively due to the significantly degraded audio quality caused by audio codec compression and channel transmission effects in real-world communication…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-12 Haohan Shi , Xiyu Shi , Safak Dogan , Saif Alzubi , Tianjin Huang , Yunxiao Zhang

With the continuous improvements of deepfake methods, forgery messages have transitioned from single-modality to multi-modal fusion, posing new challenges for existing forgery detection algorithms. In this paper, we propose AVT2-DWF, the…

Computer Vision and Pattern Recognition · Computer Science 2024-03-25 Rui Wang , Dengpan Ye , Long Tang , Yunming Zhang , Jiacheng Deng

Deep generative modeling has the potential to cause significant harm to society. Recognizing this threat, a magnitude of research into detecting so-called "Deepfakes" has emerged. This research most often focuses on the image domain, while…

Machine Learning · Computer Science 2021-11-05 Joel Frank , Lea Schönherr

Modern audio deepfake detectors built on foundation models and large training datasets achieve promising detection performance. However, they struggle with zero-day attacks, where the audio samples are generated by novel synthesis methods…

Sound · Computer Science 2026-01-12 Xuechen Liu , Xin Wang , Junichi Yamagishi

Contemporary speech enhancement predominantly relies on audio transforms that are trained to reconstruct a clean speech waveform. The development of high-performing neural network sound recognition systems has raised the possibility of…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-18 Mark R. Saddler , Andrew Francl , Jenelle Feather , Kaizhi Qian , Yang Zhang , Josh H. McDermott

Audio deepfake detection (ADD) is crucial to combat the misuse of speech synthesized from generative AI models. Existing ADD models suffer from generalization issues, with a large performance discrepancy between in-domain and out-of-domain…

Sound · Computer Science 2024-07-29 Yi Zhu , Surya Koppisetti , Trang Tran , Gaurav Bharaj

The rise of AI-driven generative models has enabled the creation of highly realistic speech deepfakes - synthetic audio signals that can imitate target speakers' voices - raising critical security concerns. Existing methods for detecting…

Sound · Computer Science 2025-03-25 Emma Coletta , Davide Salvi , Viola Negroni , Daniele Ugo Leonzio , Paolo Bestagini

Audio deepfakes are increasingly in-differentiable from organic speech, often fooling both authentication systems and human listeners. While many techniques use low-level audio features or optimization black-box model training, focusing on…

Sound · Computer Science 2025-02-21 Kevin Warren , Daniel Olszewski , Seth Layton , Kevin Butler , Carrie Gates , Patrick Traynor

In this paper, we present our comprehensive study aimed at enhancing the generalization capabilities of audio deepfake detection models. We investigate the performance of various pre-trained backbones, including Wav2Vec2, WavLM, and…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-03 Jose A. Lopez , Georg Stemmer , Héctor Cordourier Maruri

Speaker verification systems have been used in many production scenarios in recent years. Unfortunately, they are still highly prone to different kinds of spoofing attacks such as voice conversion and speech synthesis, etc. In this paper,…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-07 Junxiao Xue , Hao Zhou , Yabo Wang

Multimodal deepfakes can exhibit subtle visual artifacts and cross-modal inconsistencies, which remain challenging to detect, especially when detectors are trained primarily on curated synthetic forgeries. Such synthetic dependence can…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Sahibzada Adil Shahzad , Ammarah Hashmi , Junichi Yamagishi , Yusuke Yasuda , Yu Tsao , Chia-Wen Lin , Yan-Tsung Peng , Hsin-Min Wang

The performance of spoofing countermeasure systems depends fundamentally upon the use of sufficiently representative training data. With this usually being limited, current solutions typically lack generalisation to attacks encountered in…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-01 Hemlata Tak , Massimiliano Todisco , Xin Wang , Jee-weon Jung , Junichi Yamagishi , Nicholas Evans

Deepfake audio poses a rising threat in communication platforms, necessitating real-time detection for audio stream integrity. Unlike traditional non-real-time approaches, this study assesses the viability of employing static deepfake audio…

Automated deception detection is crucial for assisting humans in accurately assessing truthfulness and identifying deceptive behavior. Conventional contact-based techniques, like polygraph devices, rely on physiological signals to determine…

Deepfake detectors face growing challenges in generalization as new image synthesis techniques emerge. In particular, deepfakes generated by diffusion models are highly photorealistic and often evade detectors trained on GAN-based…

Computer Vision and Pattern Recognition · Computer Science 2026-04-17 Hongyuan Qi , Wenjin Hou , Hehe Fan , Jun Xiao

Audio deepfake detection systems are increasingly deployed in high-stakes security applications, yet their fairness across demographic groups remains critically underexamined. Prior work measures gender disparity but does not investigate…

Sound · Computer Science 2026-05-12 Aishwarya Fursule , Shruti Kshirsagar , Anderson R. Avila

In this paper, we perform an in-depth study of how data augmentation techniques improve synthetic or spoofed audio detection. Specifically, we propose methods to deal with channel variability, different audio compressions, different…

Sound · Computer Science 2021-10-22 Ariel Cohen , Inbal Rimon , Eran Aflalo , Haim Permuter