English
Related papers

Related papers: Phase-Aware Spoof Speech Detection Based on Res2Ne…

200 papers

Modern autoregressive speech synthesis models leveraging language models have demonstrated remarkable performance. However, the sequential nature of next token prediction in these models leads to significant latency, hindering their…

Sound · Computer Science 2025-06-04 Zijian Lin , Yang Zhang , Yougen Yuan , Yuming Yan , Jinjiang Liu , Zhiyong Wu , Pengfei Hu , Qun Yu

This paper presents a macroscopic approach to automatic detection of speech sound disorder (SSD) in child speech. Typically, SSD is manifested by persistent articulation and phonological errors on specific phonemes in the language. The…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-30 Si-Ioi Ng , Cymie Wing-Yee Ng , Jiarui Wang , Tan Lee

End-to-end approaches to anti-spoofing, especially those which operate directly upon the raw signal, are starting to be competitive with their more traditional counterparts. Until recently, all such approaches consider only the learning of…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-07 Wanying Ge , Jose Patino , Massimiliano Todisco , Nicholas Evans

The performance of automatic speaker verification (ASV) systems could be degraded by voice spoofing attacks. Most existing works aimed to develop standalone spoofing countermeasure (CM) systems. Relatively little work targeted at developing…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-05 You Zhang , Ge Zhu , Zhiyao Duan

Recent advances in sophisticated synthetic speech generated from text-to-speech (TTS) or voice conversion (VC) systems cause threats to the existing automatic speaker verification (ASV) systems. Since such synthetic speech is generated from…

Audio and Speech Processing · Electrical Eng. & Systems 2022-12-15 Youngsik Eom , Yeonghyeon Lee , Ji Sub Um , Hoirin Kim

A good training set for speech spoofing countermeasures requires diverse TTS and VC spoofing attacks, but generating TTS and VC spoofed trials for a target speaker may be technically demanding. Instead of using full-fledged TTS and VC…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-23 Xin Wang , Junichi Yamagishi

In this paper, we propose a replay attack spoofing detection system for automatic speaker verification using multitask learning of noise classes. We define the noise that is caused by the replay attack as replay noise. We explore the…

Audio and Speech Processing · Electrical Eng. & Systems 2018-10-26 Hye-Jin Shim , Jee-weon Jung , Hee-Soo Heo , Sunghyun Yoon , Ha-Jin Yu

Modern text-to-speech (TTS) and voice conversion (VC) systems produce natural sounding speech that questions the security of automatic speaker verification (ASV). This makes detection of such synthetic speech very important to safeguard ASV…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-22 Zhenzong Wu , Rohan Kumar Das , Jichen Yang , Haizhou Li

This paper describes the UZH-CL system submitted to the SASV section of the WildSpoof 2026 challenge. The challenge focuses on the integrated defense against generative spoofing attacks by requiring the simultaneous verification of speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-27 Aref Farhadipour , Ming Jin , Valeriia Vyshnevetska , Xiyang Li , Elisa Pellegrino , Srikanth Madikeri

State-of-the-art statistical parametric speech synthesis (SPSS) generally uses a vocoder to represent speech signals and parameterize them into features for subsequent modeling. Magnitude spectrum has been a dominant feature over the years.…

Sound · Computer Science 2015-10-08 Bo Fan , Siu Wa Lee , Xiaohai Tian , Lei Xie , Minghui Dong

Current audio deepfake detection has achieved remarkable performance using diverse deep learning architectures such as ResNet, and has seen further improvements with the introduction of large models (LMs) like Wav2Vec. The success of large…

Sound · Computer Science 2026-03-27 Yupei Li , Shuaijie Shao , Manuel Milling , Björn Schuller

Existing speaker verification (SV) systems often suffer from performance degradation if there is any language mismatch between model training, speaker enrollment, and test. A major cause of this degradation is that most existing SV methods…

Sound · Computer Science 2017-06-27 Lantian Li , Dong Wang , Askar Rozi , Thomas Fang Zheng

We propose supervised systems for speech activity detection (SAD) and speaker identification (SID) tasks in Fearless Steps Challenge Phase-2. The proposed systems for both the tasks share a common convolutional neural network (CNN)…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-11 Karthik Pandia D S , Cosimo Spera

In this paper, we aim to address the problem of channel robustness in speech countermeasure (CM) systems, which are used to distinguish synthetic speech from human natural speech. On the basis of two hypotheses, we suggest an approach for…

Sound · Computer Science 2023-10-10 Yongyi Zang , You Zhang , Zhiyao Duan

In this work, we address fusion of heterogeneous sensor data using wavelet-based summaries of fused self-similarity information from each sensor. The technique we develop is quite general, does not require domain specific knowledge or…

Computer Vision and Pattern Recognition · Computer Science 2019-01-08 Christopher J. Tralie , Paul Bendich , John Harer

Audio recorded in real-world environments often contains a mixture of foreground speech and background environmental sounds. With rapid advances in text-to-speech, voice conversion, and other generation models, either component can now be…

Sound · Computer Science 2026-02-06 Xueping Zhang , Han Yin , Yang Xiao , Lin Zhang , Ting Dang , Rohan Kumar Das , Ming Li

Recent research has highlighted a key issue in speech deepfake detection: models trained on one set of deepfakes perform poorly on others. The question arises: is this due to the continuously improving quality of Text-to-Speech (TTS)…

Sound · Computer Science 2024-06-13 Nicolas M. Müller , Nicholas Evans , Hemlata Tak , Philip Sperl , Konstantin Böttinger

Audio DeepFakes are utterances generated with the use of deep neural networks. They are highly misleading and pose a threat due to use in fake news, impersonation, or extortion. In this work, we focus on increasing accessibility to the…

Sound · Computer Science 2022-10-13 Piotr Kawa , Marcin Plata , Piotr Syga

Most research in synthetic speech detection (SSD) focuses on improving performance on standard noise-free datasets. However, in actual situations, noise interference is usually present, causing significant performance degradation in SSD…

Sound · Computer Science 2024-04-17 Cunhang Fan , Mingming Ding , Jianhua Tao , Ruibo Fu , Jiangyan Yi , Zhengqi Wen , Zhao Lv

Many deep learning synthetic speech generation tools are readily available. The use of synthetic speech has caused financial fraud, impersonation of people, and misinformation to spread. For this reason forensic methods that can detect…