English
Related papers

Related papers: Audio Deepfake Detection at the First Greeting: "H…

200 papers

The Audio Deep Synthesis Detection (ADD) Challenge has been held to detect generated human-like speech. With our submitted system, this paper provides an overall assessment of track 1 (Low-quality Fake Audio Detection) and track 2…

Sound · Computer Science 2022-10-12 Xiaohui Liu , Meng Liu , Lin Zhang , Linjuan Zhang , Chang Zeng , Kai Li , Nan Li , Kong Aik Lee , Longbiao Wang , Jianwu Dang

Audio deepfake detection is an emerging topic in the artificial intelligence community. The second Audio Deepfake Detection Challenge (ADD 2023) aims to spur researchers around the world to build new innovative technologies that can further…

Audio deepfake detection (ADD) is essential for preventing the misuse of synthetic voices that may infringe on personal rights and privacy. Recent zero-shot text-to-speech (TTS) models pose higher risks as they can clone voices with a…

Sound · Computer Science 2024-09-23 Yuang Li , Min Zhang , Mengxin Ren , Miaomiao Ma , Daimeng Wei , Hao Yang

Face anti-spoofing aims to discriminate the spoofing face images (e.g., printed photos) from live ones. However, adversarial examples greatly challenge its credibility, where adding some perturbation noise can easily change the predictions.…

Computer Vision and Pattern Recognition · Computer Science 2023-05-03 Songlin Yang , Wei Wang , Chenye Xu , Ziwen He , Bo Peng , Jing Dong

In this paper, Whisper, a large-scale pre-trained model for automatic speech recognition, is proposed to apply to speaker verification. A partial multi-scale feature aggregation (PMFA) approach is proposed based on a subset of Whisper…

Sound · Computer Science 2024-08-29 Yiyang Zhao , Shuai Wang , Guangzhi Sun , Zehua Chen , Chao Zhang , Mingxing Xu , Thomas Fang Zheng

This paper describes a submission to the Environment-Aware Speech and Sound Deepfake Detection Challenge (ESDD2) 2026, which addresses component-level deepfake detection using the CompSpoofV2 dataset, where speech and environmental sounds…

Sound · Computer Science 2026-05-06 Khalid Zaman , Qixuan Huang , Muhammad Uzair , Masashi Unoki

Current state-of-the-art (SOTA) codec-based audio synthesis systems can mimic anyone's voice with just a 3-second sample from that specific unseen speaker. Unfortunately, malicious attackers may exploit these technologies, causing misuse…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-12 Haibin Wu , Yuan Tseng , Hung-yi Lee

Existing fake audio detection systems perform well in in-domain testing, but still face many challenges in out-of-domain testing. This is due to the mismatch between the training and test data, as well as the poor generalizability of…

Sound · Computer Science 2023-05-24 Chenglong Wang , Jiangyan Yi , Jianhua Tao , Chuyuan Zhang , Shuai Zhang , Xun Chen

With rapid advances in audio-visual generative models, reliable forgery detection becomes increasingly critical. Existing methods for audio-visual deepfake detection typically rely on cross-modal inconsistencies. In singing, rhythmic…

Artificial Intelligence · Computer Science 2026-05-28 Ke Liu , Jiwei Wei , Wenyu Zhang , Shuchang Zhou , Ruikun Chai , Yutao Dai , Chaoning Zhang , Yang Yang

Audio deepfakes represent a growing threat to digital security and trust, leveraging advanced generative models to produce synthetic speech that closely mimics real human voices. Detecting such manipulations is especially challenging under…

Sound · Computer Science 2025-05-01 Andrea Di Pierno , Luca Guarnera , Dario Allegra , Sebastiano Battiato

The rise of AI voice-cloning technology, particularly audio Real-time Deepfakes (RTDFs), has intensified social engineering attacks by enabling real-time voice impersonation that bypasses conventional enrollment-based authentication. This…

Sound · Computer Science 2025-05-27 Govind Mittal , Arthur Jakobsson , Kelly O. Marshall , Chinmay Hegde , Nasir Memon

Deepfake speech represents a real and growing threat to systems and society. Many detectors have been created to aid in defense against speech deepfakes. While these detectors implement myriad methodologies, many rely on low-level fragments…

Deepfake audio poses a rising threat in communication platforms, necessitating real-time detection for audio stream integrity. Unlike traditional non-real-time approaches, this study assesses the viability of employing static deepfake audio…

Automatic speaker verification systems are vulnerable to a variety of access threats, prompting research into the formulation of effective spoofing detection systems to act as a gate to filter out such spoofing attacks. This study…

Sound · Computer Science 2022-11-21 Zhenyu Wang , John H. L. Hansen

Modern audio deepfake detectors built on foundation models and large training datasets achieve promising detection performance. However, they struggle with zero-day attacks, where the audio samples are generated by novel synthesis methods…

Sound · Computer Science 2026-01-12 Xuechen Liu , Xin Wang , Junichi Yamagishi

Recent advances in emotional voice conversion (EVC) have enabled the generation of expressive synthetic speech, raising new concerns in audio deepfake detection. Existing approaches treat speech as a homogeneous signal and largely overlook…

Sound · Computer Science 2026-05-06 Vamshi Nallaguntla , Shruti Kshirsagar , Anderson R. Avila

Current speech deepfake detection approaches perform satisfactorily against known adversaries; however, generalization to unseen attacks remains an open challenge. The proliferation of speech deepfakes on social media underscores the need…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-01 Ivan Kukanov , Janne Laakkonen , Tomi Kinnunen , Ville Hautamäki

Advances in AIGC technologies have enabled the synthesis of highly realistic audio deepfakes capable of deceiving human auditory perception. Although numerous audio deepfake detection (ADD) methods have been developed, most rely on local…

Sound · Computer Science 2026-02-06 Qing Wen , Haohao Li , Zhongjie Ba , Peng Cheng , Miao He , Li Lu , Kui Ren

Deepfake represents a category of face-swapping attacks that leverage machine learning models such as autoencoders or generative adversarial networks. Although the concept of the face-swapping is not new, its recent technical advances make…

Computer Vision and Pattern Recognition · Computer Science 2020-06-16 Chaofei Yang , Lei Ding , Yiran Chen , Hai Li

Automatic speaker verification (ASV) systems utilize the biometric information in human speech to verify the speaker's identity. The techniques used for performing speaker verification are often vulnerable to malicious attacks that attempt…

Sound · Computer Science 2020-11-26 Yang Gao , Jiachen Lian , Bhiksha Raj , Rita Singh