中文
相关论文

相关论文: A Data-Driven Diffusion-based Approach for Audio D…

200 篇论文

Self-supervised representations excel at many vision and speech tasks, but their potential for audio-visual deepfake detection remains underexplored. Unlike prior work that uses these features in isolation or buried within complex…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Dragos-Alexandru Boldisor , Stefan Smeu , Dan Oneata , Elisabeta Oneata

Deepfake speech detection systems are often limited to binary classification tasks and struggle to generate interpretable reasoning or provide context-rich explanations for their decisions. These models primarily extract latent embeddings…

声音 · 计算机科学 2026-04-01 Runkun Chen , Yixiong Fang , Pengyu Chang , Yuante Li , Massa Baali , Bhiksha Raj

Diffusion models have achieved remarkable success in image synthesis. However, addressing artifacts and unrealistic regions remains a critical challenge. We propose self-refining diffusion, a novel framework that enhances image generation…

计算机视觉与模式识别 · 计算机科学 2025-12-10 Seoyeon Lee , Gwangyeol Yu , Chaewon Kim , Jonghyuk Park

With the advancement of generative modeling techniques, synthetic human speech becomes increasingly indistinguishable from real, and tricky challenges are elicited for the audio deepfake detection (ADD) system. In this paper, we exploit…

声音 · 计算机科学 2024-03-05 Yujie Yang , Haochen Qin , Hang Zhou , Chengcheng Wang , Tianyu Guo , Kai Han , Yunhe Wang

Generalization in audio deepfake detection presents a significant challenge, with models trained on specific datasets often struggling to detect deepfakes generated under varying conditions and unknown algorithms. While collectively…

Deep generative models can generate high-fidelity audio conditioned on various types of representations (e.g., mel-spectrograms, Mel-frequency Cepstral Coefficients (MFCC)). Recently, such models have been used to synthesize audio waveforms…

Automatic speech recognition (ASR) is improving ever more at mimicking human speech processing. The functioning of ASR, however, remains to a large extent obfuscated by the complex structure of the deep neural networks (DNNs) they are based…

机器学习 · 计算机科学 2022-02-03 Karla Markert , Romain Parracone , Mykhailo Kulakov , Philip Sperl , Ching-Yu Kao , Konstantin Böttinger

Audio-visual temporal deepfake localization under the content-driven partial manipulation remains a highly challenging task. In this scenario, the deepfake regions are usually only spanning a few frames, with the majority of the rest…

Deep audio representation learning using multi-modal audio-visual data often leads to a better performance compared to uni-modal approaches. However, in real-world scenarios both modalities are not always available at the time of inference,…

声音 · 计算机科学 2023-02-07 Amirhossein Hajavi , Ali Etemad

The rapid advancement of deepfake technology poses a significant threat to digital media integrity. Deepfakes, synthetic media created using AI, can convincingly alter videos and audio to misrepresent reality. This creates risks of…

计算机视觉与模式识别 · 计算机科学 2024-10-21 Kashish Gandhi , Prutha Kulkarni , Taran Shah , Piyush Chaudhari , Meera Narvekar , Kranti Ghag

In this paper, we propose a deep learning based system for the task of deepfake audio detection. In particular, the draw input audio is first transformed into various spectrograms using three transformation methods of Short-time Fourier…

声音 · 计算机科学 2024-07-03 Lam Pham , Phat Lam , Truong Nguyen , Huyen Nguyen , Alexander Schindler

In this paper, we perform an in-depth study of how data augmentation techniques improve synthetic or spoofed audio detection. Specifically, we propose methods to deal with channel variability, different audio compressions, different…

声音 · 计算机科学 2021-10-22 Ariel Cohen , Inbal Rimon , Eran Aflalo , Haim Permuter

In the contemporary digital age, the proliferation of deepfakes presents a formidable challenge to the sanctity of information dissemination. Audio deepfakes, in particular, can be deceptively realistic, posing significant risks in…

声音 · 计算机科学 2024-02-28 Karthik Sivarama Krishnan , Koushik Sivarama Krishnan

Audio deepfake detection (ADD) faces critical generalization challenges due to diverse real-world spoofing attacks and domain variations. However, existing methods primarily rely on Euclidean distances, failing to adequately capture the…

Proactive watermarking offers a promising approach for deepfake tamper detection and localization in short-form videos. However, existing methods often decouple audio and visual evidence and assume that watermark signals remain reliable…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Bokang Zeng , Zheng Gao , Xiaoyu Li , Xiaoyan Feng , Jiaojiao Jiang

Recent advances in technology for hyper-realistic visual and audio effects provoke the concern that deepfake videos of political speeches will soon be indistinguishable from authentic video recordings. The conventional wisdom in…

人机交互 · 计算机科学 2024-01-17 Matthew Groh , Aruna Sankaranarayanan , Nikhil Singh , Dong Young Kim , Andrew Lippman , Rosalind Picard

Building upon Diff-A-Riff, a latent diffusion model for musical instrument accompaniment generation, we present a series of improvements targeting quality, diversity, inference speed, and text-driven control. First, we upgrade the…

声音 · 计算机科学 2024-10-31 Javier Nistal , Marco Pasini , Stefan Lattner

Current DeepFake detection scenarios are mostly binary, yet data manipulation can vary across audio, video, or both, whose variability is not captured in binary settings. Four-class audio-visual formulations address this by discriminating…

计算机视觉与模式识别 · 计算机科学 2026-05-01 Sharayu Nilesh Deshmukh , Kailash A. Hambarde , Joana C. Costa , Hugo Proença , Tiago Roxo

Audio DeepFakes are utterances generated with the use of deep neural networks. They are highly misleading and pose a threat due to use in fake news, impersonation, or extortion. In this work, we focus on increasing accessibility to the…

声音 · 计算机科学 2022-10-13 Piotr Kawa , Marcin Plata , Piotr Syga

In the field of deepfake detection, previous studies focus on using reconstruction or mask and prediction methods to train pre-trained models, which are then transferred to fake audio detection training where the encoder is used to extract…