English
Related papers

Related papers: ReMASC: Realistic Replay Attack Corpus for Voice C…

200 papers

Voice authentication has become an integral part in security-critical operations, such as bank transactions and call center conversations. The vulnerability of automatic speaker verification systems (ASVs) to spoofing attacks instigated the…

Cryptography and Security · Computer Science 2021-08-02 Andre Kassis , Urs Hengartner

We show how replay attacks undermine audio deepfake detection: By playing and re-recording deepfake audio through various speakers and microphones, we make spoofed samples appear authentic to the detection model. To study this phenomenon in…

Constructing a dataset for replay spoofing detection requires a physical process of playing an utterance and re-recording it, presenting a challenge to the collection of large-scale datasets. In this study, we propose a self-supervised…

Machine Learning · Computer Science 2020-08-20 Hye-jin Shim , Hee-Soo Heo , Jee-weon Jung , Ha-Jin Yu

The rapid advancements in AI voice cloning, fueled by machine learning, have significantly impacted text-to-speech (TTS) and voice conversion (VC) fields. While these developments have led to notable progress, they have also raised concerns…

Sound · Computer Science 2025-02-17 Qingyuan Fei , Wenjie Hou , Xuan Hai , Xin Liu

The ASVspoof initiative was conceived to spearhead research in anti-spoofing for automatic speaker verification (ASV). This paper describes the third in a series of bi-annual challenges: ASVspoof 2019. With the challenge database and…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-23 Andreas Nautsch , Xin Wang , Nicholas Evans , Tomi Kinnunen , Ville Vestman , Massimiliano Todisco , Héctor Delgado , Md Sahidullah , Junichi Yamagishi , Kong Aik Lee

Over the last few years, a rapidly increasing number of Internet-of-Things (IoT) systems that adopt voice as the primary user input have emerged. These systems have been shown to be vulnerable to various types of voice spoofing attacks.…

Cryptography and Security · Computer Science 2018-03-28 Yuan Gong , Christian Poellabauer

Voice anonymization systems aim to protect speaker privacy by obscuring vocal traits while preserving the linguistic content relevant for downstream applications. However, because these linguistic cues remain intact, they can be exploited…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-21 Ahmad Aloradi , Ünal Ege Gaznepoglu , Emanuël A. P. Habets , Daniel Tenbrinck

Generative audio technologies now enable highly realistic voice cloning and real-time voice conversion, increasing the risk of impersonation, fraud, and misinformation in communication channels such as phone and video calls. This study…

Sound · Computer Science 2026-01-09 Prajwal Chinchmalatpure , Suyash Chinchmalatpure , Siddharth Chavan

Deep Learning has advanced Automatic Speaker Verification (ASV) in the past few years. Although it is known that deep learning-based ASV systems are vulnerable to adversarial examples in digital access, there are few studies on adversarial…

Sound · Computer Science 2024-01-04 Jiaqi Li , Li Wang , Liumeng Xue , Lei Wang , Zhizheng Wu

Automatic speaker verification (ASV) plays a critical role in security-sensitive environments. Regrettably, the reliability of ASV has been undermined by the emergence of spoofing attacks, such as replay and synthetic speech, as well as…

Sound · Computer Science 2023-06-27 Haibin Wu , Jiawen Kang , Lingwei Meng , Helen Meng , Hung-yi Lee

Diffusion Models (DMs) have achieved remarkable success in realistic voice cloning (VC), while they also increase the risk of malicious misuse. Existing proactive defenses designed for traditional VC models aim to disrupt the forgery…

Sound · Computer Science 2025-12-10 Qianyue Hu , Junyan Wu , Wei Lu , Xiangyang Luo

As the use of Voice Processing Systems (VPS) continues to become more prevalent in our daily lives through the increased reliance on applications such as commercial voice recognition devices as well as major text-to-speech software, the…

Cryptography and Security · Computer Science 2021-12-28 Robert Chang , Logan Kuo , Arthur Liu , Nader Sehatbakhsh

Voice Recognition Systems (VRSs) employ deep learning for speech recognition and speaker recognition. They have been widely deployed in various real-world applications, from intelligent voice assistance to telephony surveillance and…

Cryptography and Security · Computer Science 2023-07-26 Baochen Yan , Jiahe Lan , Zheng Yan

The task of deepfakes detection is far from being solved by speech or vision researchers. Several publicly available databases of fake synthetic video and speech were built to aid the development of detection methods. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2023-11-30 Pavel Korshunov , Haolin Chen , Philip N. Garner , Sebastien Marcel

In this paper, we propose a replay attack spoofing detection system for automatic speaker verification using multitask learning of noise classes. We define the noise that is caused by the replay attack as replay noise. We explore the…

Audio and Speech Processing · Electrical Eng. & Systems 2018-10-26 Hye-Jin Shim , Jee-weon Jung , Hee-Soo Heo , Sunghyun Yoon , Ha-Jin Yu

Despite the recent advancements in speech recognition, there are still difficulties in accurately transcribing conversational and emotional speech in noisy and reverberant acoustic environments. This poses a particular challenge in the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-26 Sangeet Sagar , Mirco Ravanelli , Bernd Kiefer , Ivana Kruijff Korbayova , Josef van Genabith

Describes an audio dataset of spoken words designed to help train and evaluate keyword spotting systems. Discusses why this task is an interesting challenge, and why it requires a specialized dataset that is different from conventional…

Computation and Language · Computer Science 2018-04-11 Pete Warden

Modern speaker recognition system relies on abundant and balanced datasets for classification training. However, diverse defective datasets, such as partially-labelled, small-scale, and imbalanced datasets, are common in real-world…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-03 Ruijie Tao , Zhan Shi , Yidi Jiang , Tianchi Liu , Haizhou Li

In this paper, we present an acoustic database, designed to drive and support research on voiced enabled technologies inside moving vehicles. The recording process involves (i) recordings of acoustic impulse responses, acquired under static…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-28 Nikolaos Stefanakis , Marinos Kalaitzakis , Andreas Symiakakis , Stefanos Papadakis , Despoina Pavlidi

The widespread adoption of voice-activated systems has modified routine human-machine interaction but has also introduced new vulnerabilities. This paper investigates the susceptibility of automatic speech recognition (ASR) algorithms in…

Cryptography and Security · Computer Science 2024-04-09 Forrest McKee , David Noever