中文
相关论文

相关论文: Unmasking real-world audio deepfakes: A data-centr…

200 篇论文

Recent research has highlighted a key issue in speech deepfake detection: models trained on one set of deepfakes perform poorly on others. The question arises: is this due to the continuously improving quality of Text-to-Speech (TTS)…

声音 · 计算机科学 2024-06-13 Nicolas M. Müller , Nicholas Evans , Hemlata Tak , Philip Sperl , Konstantin Böttinger

The rapid advancement of speech generation technology has led to the widespread proliferation of deepfake speech across social media platforms. While deepfake audio countermeasures (CMs) achieve promising results on public datasets, their…

声音 · 计算机科学 2025-08-15 Yuankun Xie , Ruibo Fu , Xiaopeng Wang , Zhiyong Wang , Ya Li , Zhengqi Wen , Haonnan Cheng , Long Ye

Generative AI advances rapidly, allowing the creation of very realistic manipulated video and audio. This progress presents a significant security and ethical threat, as malicious users can exploit DeepFake techniques to spread…

多媒体 · 计算机科学 2025-06-09 Marcel Klemt , Carlotta Segna , Anna Rohrbach

With the proliferation of deepfake audio, there is an urgent need to investigate their attribution. Current source tracing methods can effectively distinguish in-distribution (ID) categories. However, the rapid evolution of deepfake…

声音 · 计算机科学 2024-06-11 Yuankun Xie , Ruibo Fu , Zhengqi Wen , Zhiyong Wang , Xiaopeng Wang , Haonnan Cheng , Long Ye , Jianhua Tao

Audio deepfake detection (ADD) is essential for preventing the misuse of synthetic voices that may infringe on personal rights and privacy. Recent zero-shot text-to-speech (TTS) models pose higher risks as they can clone voices with a…

声音 · 计算机科学 2024-09-23 Yuang Li , Min Zhang , Mengxin Ren , Miaomiao Ma , Daimeng Wei , Hao Yang

Existing deepfake detection research has primarily focused on scenarios where the manipulated subject is actively speaking, i.e., generating fabricated content by altering the speaker's appearance or voice. However, in realistic interaction…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Miao Liu , Fangda Wei , Jing Wang , Xinyuan Qian

Audio deepfakes generated by modern TTS and voice conversion systems are increasingly difficult to distinguish from real speech, raising serious risks for security and online trust. While state-of-the-art self-supervised models provide rich…

声音 · 计算机科学 2026-02-03 Zhili Nicholas Liang , Soyeon Caren Han , Qizhou Wang , Christopher Leckie

The rapid development of audio-driven talking head generators and advanced Text-To-Speech (TTS) models has led to more sophisticated temporal deepfakes. These advances highlight the need for robust methods capable of detecting and…

音频与语音处理 · 电气工程与系统科学 2025-08-12 Ivan Kukanov , Jun Wah Ng

The growing sophistication of speech generated by Artificial Intelligence (AI) has introduced new challenges in audio deepfake detection. Text-to-speech (TTS) and voice conversion (VC) technologies can create highly convincing synthetic…

声音 · 计算机科学 2026-03-17 Vamshi Nallaguntla , Aishwarya Fursule , Shruti Kshirsagar , Anderson R. Avila

In this paper, we perform an in-depth study of how data augmentation techniques improve synthetic or spoofed audio detection. Specifically, we propose methods to deal with channel variability, different audio compressions, different…

声音 · 计算机科学 2021-10-22 Ariel Cohen , Inbal Rimon , Eran Aflalo , Haim Permuter

Significant advances in deep learning have obtained hallmark accuracy rates for various computer vision applications. However, advances in deep generative models have also led to the generation of very realistic fake content, also known as…

计算机视觉与模式识别 · 计算机科学 2021-10-06 Sreeraj Ramachandran , Aakash Varma Nadimpalli , Ajita Rattani

The proliferation of malicious deepfake applications has ignited substantial public apprehension, casting a shadow of doubt upon the integrity of digital media. Despite the development of proficient deepfake detection mechanisms, they…

密码学与安全 · 计算机科学 2024-03-12 Hong Sun , Ziqiang Li , Lei Liu , Bin Li

The increasing realism and accessibility of deepfakes have raised critical concerns about media authenticity and information integrity. Despite recent advances, deepfake detection models often struggle to generalize beyond their training…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Stelios Mylonas , Symeon Papadopoulos

Deepfakes, leveraging advanced AIGC (Artificial Intelligence-Generated Content) techniques, create hyper-realistic synthetic images and videos of human faces, posing a significant threat to the authenticity of social media. While this…

计算机视觉与模式识别 · 计算机科学 2025-10-10 Junyu Shi , Minghui Li , Junguo Zuo , Zhifei Yu , Yipeng Lin , Shengshan Hu , Ziqi Zhou , Yechao Zhang , Wei Wan , Yinzhe Xu , Leo Yu Zhang

In this paper, we present our comprehensive study aimed at enhancing the generalization capabilities of audio deepfake detection models. We investigate the performance of various pre-trained backbones, including Wav2Vec2, WavLM, and…

音频与语音处理 · 电气工程与系统科学 2025-07-03 Jose A. Lopez , Georg Stemmer , Héctor Cordourier Maruri

Deepfake is content or material that is synthetically generated or manipulated using artificial intelligence (AI) methods, to be passed off as real and can include audio, video, image, and text synthesis. This survey has been conducted with…

声音 · 计算机科学 2021-11-30 Zahra Khanjani , Gabrielle Watson , Vandana P. Janeja

To train transcriptor models that produce robust results, a large and diverse labeled dataset is required. Finding such data with the necessary characteristics is a challenging task, especially for languages less popular than English.…

声音 · 计算机科学 2026-05-01 Alexandre R. Ferreira , Cláudio E. C. Campelo

Good datasets are essential for developing and benchmarking any machine learning system. Their importance is even more extreme for safety critical applications such as deepfake detection - the focus of this paper. Here we reveal that two of…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Stefan Smeu , Dragos-Alexandru Boldisor , Dan Oneata , Elisabeta Oneata

Many datasets have been designed to further the development of fake audio detection. However, fake utterances in previous datasets are mostly generated by altering timbre, prosody, linguistic content or channel noise of original audio.…

This perspective calls for scholars across disciplines to address the challenge of audio deepfake detection and discernment through an interdisciplinary lens across Artificial Intelligence methods and linguistics. With an avalanche of tools…

声音 · 计算机科学 2024-11-12 Vandana P. Janeja , Christine Mallinson