English
Related papers

Related papers: MultiAPI Spoof: A Multi-API Dataset and Local-Atte…

200 papers

This study investigates the explainability of embedding representations, specifically those used in modern audio spoofing detection systems based on deep neural networks, known as spoof embeddings. Building on established work in speaker…

Sound · Computer Science 2024-12-25 Xuechen Liu , Junichi Yamagishi , Md Sahidullah , Tomi kinnunen

Speech enhancement (SE) aims to extract the clean waveform from noise-contaminated measurements to improve the speech quality and intelligibility. Although learning-based methods can perform much better than traditional counterparts, the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-23 Haoyin Yan , Jie Zhang , Cunhang Fan , Yeping Zhou , Peiqi Liu

With recent advances in speech synthesis, synthetic data is becoming a viable alternative to real data for training speech recognition models. However, machine learning with synthetic data is not trivial due to the gap between the synthetic…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-25 Ting-Yao Hu , Mohammadreza Armandpour , Ashish Shrivastava , Jen-Hao Rick Chang , Hema Koppula , Oncel Tuzel

The rhythm of bonafide speech is often difficult to replicate, which causes that the fundamental frequency (F0) of synthetic speech is significantly different from that of real speech. It is expected that the F0 feature contains the…

Sound · Computer Science 2024-07-09 Cunhang Fan , Jun Xue , Jianhua Tao , Jiangyan Yi , Chenglong Wang , Chengshi Zheng , Zhao Lv

Self-supervised learning (SSL) has transformed speech processing, with benchmarks such as SUPERB establishing fair comparisons across diverse downstream tasks. Despite it's security-critical importance, Audio deepfake detection has remained…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-10 Hashim Ali , Nithin Sai Adupa , Surya Subramani , Hafiz Malik

Research in the past several years has boosted the performance of automatic speaker verification systems and countermeasure systems to deliver low Equal Error Rates (EERs) on each system. However, research on joint optimization of both…

Sound · Computer Science 2022-03-28 Zhongwei Teng , Quchen Fu , Jules White , Maria E. Powell , Douglas C. Schmidt

Current audio deepfake detection has achieved remarkable performance using diverse deep learning architectures such as ResNet, and has seen further improvements with the introduction of large models (LMs) like Wav2Vec. The success of large…

Sound · Computer Science 2026-03-27 Yupei Li , Shuaijie Shao , Manuel Milling , Björn Schuller

Detecting partial deepfake speech is challenging because manipulations occur only in short regions while the surrounding audio remains authentic. However, existing detection methods are fundamentally limited by the quality of available…

Sound · Computer Science 2025-12-16 Menglu Li , Majd Alber , Ramtin Asgarianamiri , Lian Zhao , Xiao-Ping Zhang

In recent years, speech processing algorithms have seen tremendous progress primarily due to the deep learning renaissance. This is especially true for speech separation where the time-domain audio separation network (TasNet) has led to…

Sound · Computer Science 2021-03-30 Morten Kolbæk , Zheng-Hua Tan , Søren Holdt Jensen , Jesper Jensen

Speech separation seeks to separate individual speech signals from a speech mixture. Typically, most separation models are trained on synthetic data due to the unavailability of target reference in real-world cocktail party scenarios. As a…

Sound · Computer Science 2024-11-06 Wupeng Wang , Zexu Pan , Xinke Li , Shuai Wang , Haizhou Li

All existing databases of spoofed speech contain attack data that is spoofed in its entirety. In practice, it is entirely plausible that successful attacks can be mounted with utterances that are only partially spoofed. By definition,…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-16 Lin Zhang , Xin Wang , Erica Cooper , Junichi Yamagishi , Jose Patino , Nicholas Evans

ASVspoof 5 is the fifth edition in a series of challenges which promote the study of speech spoofing and deepfake detection solutions. A significant change from previous challenge editions is a new crowdsourced database collected from a…

This paper describes the UZH-CL system submitted to the SASV section of the WildSpoof 2026 challenge. The challenge focuses on the integrated defense against generative spoofing attacks by requiring the simultaneous verification of speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-27 Aref Farhadipour , Ming Jin , Valeriia Vyshnevetska , Xiyang Li , Elisa Pellegrino , Srikanth Madikeri

This paper presents the DFKI-Speech system developed for the WildSpoof Challenge under the Spoofing aware Automatic Speaker Verification (SASV) track. We propose a robust SASV framework in which a spoofing detector and a speaker…

Now-a-days, speech-based biometric systems such as automatic speaker verification (ASV) are highly prone to spoofing attacks by an imposture. With recent development in various voice conversion (VC) and speech synthesis (SS) algorithms,…

Sound · Computer Science 2016-11-18 Dipjyoti Paul , Monisankha Pal , Goutam Saha

Current state-of-the-art (SOTA) codec-based audio synthesis systems can mimic anyone's voice with just a 3-second sample from that specific unseen speaker. Unfortunately, malicious attackers may exploit these technologies, causing misuse…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-12 Haibin Wu , Yuan Tseng , Hung-yi Lee

The recent developments in technology have re-warded us with amazing audio synthesis models like TACOTRON and WAVENETS. On the other side, it poses greater threats such as speech clones and deep fakes, that may go undetected. To tackle…

Machine Learning · Computer Science 2021-07-27 Arun Kumar Singh , Priyanka Singh , Karan Nathwani

Face anti-spoofing is crucial for ensuring the security and reliability of face recognition systems. Several existing face anti-spoofing methods utilize GAN-like networks to detect presentation attacks by estimating the noise pattern of a…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Bin Zhang , Xiangyu Zhu , Xiaoyu Zhang , Zhen Lei

Paralinguistic sounds, like laughter and sighs, are crucial for synthesizing more realistic and engaging speech. However, existing methods typically depend on proprietary datasets, while publicly available resources often suffer from…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-30 Bingsong Bai , Qihang Lu , Wenbing Yang , Zihan Sun , Yueran Hou , Peilei Jia , Songbai Pu , Ruibo Fu , Yingming Gao , Ya Li , Jun Gao

This study focuses on building effective spoofing countermeasures (CMs) for non-native speech, specifically targeting Indonesian and Thai speakers. We constructed a dataset comprising both native and non-native speech to facilitate our…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-03 Aulia Adila , Candy Olivia Mawalim , Masashi Unoki