English
Related papers

Related papers: A Data-Driven Diffusion-based Approach for Audio D…

200 papers

The field of eXplainable Artificial Intelligence (XAI) has greatly advanced in recent years, but progress has mainly been made in computer vision and natural language processing. For time series, where the input is often not interpretable,…

Machine Learning · Computer Science 2023-03-14 Johanna Vielhaben , Sebastian Lapuschkin , Grégoire Montavon , Wojciech Samek

Voice authentication systems deployed at the network edge face dual threats: a) sophisticated deepfake synthesis attacks and b) control-plane poisoning in distributed federated learning protocols. We present a framework coupling…

Sound · Computer Science 2025-12-09 Alireza Mohammadi , Keshav Sood , Dhananjay Thiruvady , Asef Nazari

Aliasing artifacts in renderings produced by Neural Radiance Field (NeRF) is a long-standing but complex issue in the field of 3D implicit representation, which arises from a multitude of intricate causes and was mitigated by designing more…

Computer Vision and Pattern Recognition · Computer Science 2024-07-11 Ganlin Yang , Kaidong Zhang , Jingjing Fu , Dong Liu

This paper presents the Deep learning-based Perceptual Audio Quality metric (DeePAQ) for evaluating general audio quality. Our approach leverages metric learning together with the music foundation model MERT, guided by surrogate labels, to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-15 Guanxin Jiang , Andreas Brendel , Pablo M. Delgado , Jürgen Herre

In this paper, we present our comprehensive study aimed at enhancing the generalization capabilities of audio deepfake detection models. We investigate the performance of various pre-trained backbones, including Wav2Vec2, WavLM, and…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-03 Jose A. Lopez , Georg Stemmer , Héctor Cordourier Maruri

The rapid evolution of speech synthesis and voice conversion has raised substantial concerns due to the potential misuse of such technology, prompting a pressing need for effective audio deepfake detection mechanisms. Existing detection…

Sound · Computer Science 2023-12-18 Xiaohui Zhang , Jiangyan Yi , Chenglong Wang , Chuyuan Zhang , Siding Zeng , Jianhua Tao

Deep learning classifiers are prone to latching onto dominant confounders present in a dataset rather than on the causal markers associated with the target class, leading to poor generalization and biased predictions. Although…

Computer Vision and Pattern Recognition · Computer Science 2024-05-16 Nima Fathi , Amar Kumar , Brennan Nichyporuk , Mohammad Havaei , Tal Arbel

Diffusion-based speech enhancement (SE) achieves natural-sounding speech and strong generalization, yet suffers from key limitations like generative artifacts and high inference latency. In this work, we systematically study artifact…

Audio deepfake detection has recently garnered public concern due to its implications for security and reliability. Traditional deep learning methods have been widely applied to this task but often lack generalisability when confronted with…

Sound · Computer Science 2025-12-16 Yupei Li , Li Wang , Yuxiang Wang , Lei Wang , Rizhao Cai , Jie Shi , Björn W. Schuller , Zhizheng Wu

Trust and credibility in machine learning models is bolstered by the ability of a model to explain itsdecisions. While explainability of deep learning models is a well-known challenge, a further chal-lenge is clarity of the explanation…

Machine Learning · Computer Science 2020-11-30 hsan Ullah , Andre Rios , Vaibhav Gala , Susan Mckeever

With the advancement of audio generation, generative models can produce highly realistic audios. However, the proliferation of deepfake general audio can pose negative consequences. Therefore, we propose a new task, deepfake general audio…

Sound · Computer Science 2024-06-13 Zeyu Xie , Baihan Li , Xuenan Xu , Zheng Liang , Kai Yu , Mengyue Wu

Substantial progress in spoofing and deepfake detection has been made in recent years. Nonetheless, the community has yet to make notable inroads in providing an explanation for how a classifier produces its output. The dominance of black…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-29 Wanying Ge , Jose Patino , Massimiliano Todisco , Nicholas Evans

Depth estimation aims to predict dense depth maps. In autonomous driving scenes, sparsity of annotations makes the task challenging. Supervised models produce concave objects due to insufficient structural information. They overfit to valid…

Computer Vision and Pattern Recognition · Computer Science 2023-08-07 Jiaqi Li , Yiran Wang , Zihao Huang , Jinghong Zheng , Ke Xian , Zhiguo Cao , Jianming Zhang

Audio deepfake detection (ADD) is essential for preventing the misuse of synthetic voices that may infringe on personal rights and privacy. Recent zero-shot text-to-speech (TTS) models pose higher risks as they can clone voices with a…

Sound · Computer Science 2024-09-23 Yuang Li , Min Zhang , Mengxin Ren , Miaomiao Ma , Daimeng Wei , Hao Yang

The growing prevalence of real-world deepfakes presents a critical challenge for existing detection systems, which are often evaluated on datasets collected just for scientific purposes. To address this gap, we introduce a novel dataset of…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-30 David Combei , Adriana Stan , Dan Oneata , Nicolas Müller , Horia Cucu

Digital pathology plays a vital role across modern medicine, offering critical insights for disease diagnosis, prognosis, and treatment. However, histopathology images often contain artifacts introduced during slide preparation and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Zhenzhen Wang , Zhongliang Zhou , Zhuoyu Wen , Jeong Hwan Kook , John B Wojcik , John Kang

Digital audio tampering detection can be used to verify the authenticity of digital audio. However, most current methods use standard electronic network frequency (ENF) databases for visual comparison analysis of ENF continuity of digital…

Sound · Computer Science 2022-10-20 Zhifeng Wang , Yao Yang , Chunyan Zeng , Shuai Kong , Shixiong Feng , Nan Zhao

Audio deepfake detection systems trained on one dataset often fail when deployed on data from different sources due to distributional shifts in recording conditions, synthesis methods, and acoustic environments. We present a modular…

Sound · Computer Science 2026-03-10 Urawee Thani , Gagandeep Singh , Priyanka Singh

The rapid development of deep learning and generative AI technologies has profoundly transformed the digital contact landscape, creating realistic Deepfake that poses substantial challenges to public trust and digital media integrity. This…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Ying Xu , Marius Pedersen , Kiran Raja

Speech deepfake detection (SDD) focuses on identifying whether a given speech signal is genuine or has been synthetically generated. Existing audio large language model (LLM)-based methods excel in content understanding; however, their…

Sound · Computer Science 2026-02-02 Xiaoxuan Guo , Yuankun Xie , Haonan Cheng , Jiayi Zhou , Jian Liu , Hengyan Huang , Long Ye , Qin Zhang