English
Related papers

Related papers: Robust Audio Anti-Spoofing with Fusion-Reconstruct…

200 papers

Speech emotion recognition systems (SER) can achieve high accuracy when the training and test data are identically distributed, but this assumption is frequently violated in practice and the performance of SER systems plummet against…

Sound · Computer Science 2020-07-28 Siddique Latif , Rajib Rana , Sara Khalifa , Raja Jurdak , Björn W. Schuller

This paper presents the SZU-AFS anti-spoofing system, designed for Track 1 of the ASVspoof 5 Challenge under open conditions. The system is built with four stages: selecting a baseline model, exploring effective data augmentation (DA)…

Sound · Computer Science 2024-08-20 Yuxiong Xu , Jiafeng Zhong , Sengui Zheng , Zefeng Liu , Bin Li

The Automatic Speaker Verification systems have potential in biometrics applications for logical control access and authentication. A lot of things happen to be at stake if the ASV system is compromised. The preliminary work presents a…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-28 Rohit Arora

Recent works on speech spoofing countermeasures still lack generalization ability to unseen spoofing attacks. This is one of the key issues of ASVspoof challenges especially with the rapid development of diverse and high-quality spoofing…

Audio and Speech Processing · Electrical Eng. & Systems 2022-01-25 Monisankha Pal , Aditya Raikar , Ashish Panda , Sunil Kumar Kopparapu

In this paper, we investigate the impact of different standard environmental sound representations (spectrograms) on the recognition performance and adversarial attack robustness of a victim residual convolutional neural network. Averaged…

Audio and Speech Processing · Electrical Eng. & Systems 2021-01-19 Mohammad Esmaeilpour , Patrick Cardinal , Alessandro Lameiras Koerich

Recent advancements in multi-scale architectures have demonstrated exceptional performance in image denoising tasks. However, existing architectures mainly depends on a fixed single-input single-output Unet architecture, ignoring the…

Computer Vision and Pattern Recognition · Computer Science 2025-06-23 Xu Zhao , Chen Zhao , Xiantao Hu , Hongliang Zhang , Ying Tai , Jian Yang

Deepfake (DF) audio detectors still struggle to generalize to out of distribution inputs. A central reason is spectral bias, the tendency of neural networks to learn low-frequency structure before high-frequency (HF) details, which both…

Sound · Computer Science 2025-11-27 Ido Nitzan HIdekel , Gal lifshitz , Khen Cohen , Dan Raviv

Motivated by the attention mechanism of the human visual system and recent developments in the field of machine translation, we introduce our attention-based and recurrent sequence to sequence autoencoders for fully unsupervised…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-20 Shahin Amiriparian , Pawel Winokurow , Vincent Karas , Sandra Ottl , Maurice Gerczuk , Björn W. Schuller

While deepfake speech detectors built on large self-supervised learning (SSL) models achieve high accuracy, employing standard ensemble fusion to further enhance robustness often results in oversized systems with diminishing returns. To…

Sound · Computer Science 2026-04-03 Vojtěch Staněk , Martin Perešíni , Lukáš Sekanina , Anton Firc , Kamil Malinka

Deepfakes are AI-synthesized multimedia data that may be abused for spreading misinformation. Deepfake generation involves both visual and audio manipulation. To detect audio-visual deepfakes, previous studies commonly employ two relatively…

Sound · Computer Science 2025-06-10 Kuiyuan Zhang , Wenjie Pei , Rushi Lan , Yifang Guo , Zhongyun Hua

The efficient fusion of depth maps is a key part of most state-of-the-art 3D reconstruction methods. Besides requiring high accuracy, these depth fusion methods need to be scalable and real-time capable. To this end, we present a novel…

Computer Vision and Pattern Recognition · Computer Science 2020-04-06 Silvan Weder , Johannes L. Schönberger , Marc Pollefeys , Martin R. Oswald

Photoacoustic (PA) imaging is a biomedical imaging modality capable of acquiring high-contrast images of optical absorption at depths much greater than traditional optical imaging techniques. However, practical instrumentation and geometry…

Computer Vision and Pattern Recognition · Computer Science 2021-06-02 Mengjie Guo , Hengrong Lan , Changchun Yang , Fei Gao

Pattern recognition from audio signals is an active research topic encompassing audio tagging, acoustic scene classification, music classification, and other areas. Spectrogram and mel-frequency cepstral coefficients (MFCC) are among the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-18 Md. Istiaq Ansari , Taufiq Hasan

Supervised contrastive learning (SupCon) is widely used to shape representations, but has seen limited targeted study for audio deepfake detection. Existing work typically combines contrastive terms with broader pipelines; however, the…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-30 Jaskirat Sudan , Hashim Ali , Surya Subramani , Hafiz Malik

Face anti-spoofing (FAS) plays a vital role in securing the face recognition systems from presentation attacks. Most existing FAS methods capture various cues (e.g., texture, depth and reflection) to distinguish the live faces from the…

Computer Vision and Pattern Recognition · Computer Science 2020-07-07 Zitong Yu , Xiaobai Li , Xuesong Niu , Jingang Shi , Guoying Zhao

Detecting manipulated media has now become a pressing issue with the recent rise of deepfakes. Most existing approaches fail to generalize across diverse datasets and generation techniques. We thus propose a novel ensemble framework,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Vrushank Ahire , Aniruddh Muley , Shivam Zample , Siddharth Verma , Pranav Menon , Surbhi Madan , Abhinav Dhall

Biomedical imaging is unequivocally dependent on the ability to reconstruct interpretable and high-quality images from acquired sensor data. This reconstruction process is pivotal across many applications, spanning from magnetic resonance…

Signal Processing · Electrical Eng. & Systems 2019-09-24 Ben Luijten , Regev Cohen , Frederik J. de Bruijn , Harold A. W. Schmeitz , Massimo Mischi , Yonina C. Eldar , Ruud J. G. van Sloun

Generalization in audio deepfake detection presents a significant challenge, with models trained on specific datasets often struggling to detect deepfakes generated under varying conditions and unknown algorithms. While collectively…

Stream fusion, also known as system combination, is a common technique in automatic speech recognition for traditional hybrid hidden Markov model approaches, yet mostly unexplored for modern deep neural network end-to-end model…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-15 Timo Lohrenz , Zhengyang Li , Tim Fingscheidt

Traditional defenses against Deep Leakage (DL) attacks in Federated Learning (FL) primarily focus on obfuscation, introducing noise, transformations or encryption to degrade an attacker's ability to reconstruct private data. While effective…

Cryptography and Security · Computer Science 2026-01-22 Isaac Baglin , Xiatian Zhu , Simon Hadfield