English
Related papers

Related papers: Pseudo-Siamese Network based Timbre-reserved Black…

200 papers

Widely used deep learning models are found to have poor robustness. Little noises can fool state-of-the-art models into making incorrect predictions. While there is a great deal of high-performance attack generation methods, most of them…

Machine Learning · Computer Science 2022-08-26 Xinyi Wang , Simon Yusuf Enoch , Dong Seong Kim

Speaker identification models are vulnerable to carefully designed adversarial perturbations of their input signals that induce misclassification. In this work, we propose a white-box steganography-inspired adversarial attack that generates…

The use of deep networks to extract embeddings for speaker recognition has proven successfully. However, such embeddings are susceptible to performance degradation due to the mismatches among the training, enrollment, and test conditions.…

Sound · Computer Science 2019-04-30 Zhong Meng , Yong Zhao , Jinyu Li , Yifan Gong

A judicious combination of dictionary learning methods, block sparsity and source recovery algorithm are used in a hierarchical manner to identify the noises and the speakers from a noisy conversation between two people. Conversations are…

Sound · Computer Science 2016-10-31 K V Vijay Girish , A G Ramakrishnan , T V Ananthapadmanabha

Thanks to the popularisation of transformer-based models, speech recognition (SR) is gaining traction in various application fields, such as industrial and robotics environments populated with mission-critical devices. While…

Cryptography and Security · Computer Science 2024-09-20 Jonatan Bartolini , Todor Stoyanov , Alberto Giaretta

In this paper, we propose PhantomSound, a query-efficient black-box attack toward voice assistants. Existing black-box adversarial attacks on voice assistants either apply substitution models or leverage the intermediate model output to…

Cryptography and Security · Computer Science 2023-09-14 Hanqing Guo , Guangjing Wang , Yuanda Wang , Bocheng Chen , Qiben Yan , Li Xiao

In this work, we propose a multi-target backdoor attack against speaker identification using position-independent clicking sounds as triggers. Unlike previous single-target approaches, our method targets up to 50 speakers simultaneously,…

Audio deepfake detection is crucial to combat the malicious use of AI-synthesized speech. Among many efforts undertaken by the community, the ASVspoof challenge has become one of the benchmarks to evaluate the generalizability and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-11 Yi Zhu , Chirag Goel , Surya Koppisetti , Trang Tran , Ankur Kumar , Gaurav Bharaj

This research project investigates the application of deep learning to timbre transfer, where the timbre of a source audio can be converted to the timbre of a target audio with minimal loss in quality. The adopted approach combines…

Sound · Computer Science 2021-10-12 Russell Sammut Bonnici , Charalampos Saitis , Martin Benning

Fooling deep neural networks with adversarial input have exposed a significant vulnerability in the current state-of-the-art systems in multiple domains. Both black-box and white-box approaches have been used to either replicate the model…

Cryptography and Security · Computer Science 2019-07-04 Shreya Khare , Rahul Aralikatte , Senthil Mani

Adversarial attacks remain a significant threat that can jeopardize the integrity of Machine Learning (ML) models. In particular, query-based black-box attacks can generate malicious noise without having access to the victim model's…

Cryptography and Security · Computer Science 2025-03-18 Jeonghwan Park , Niall McLaughlin , Ihsen Alouani

Speaker recognition systems (SRSs) have recently been shown to be vulnerable to adversarial attacks, raising significant security concerns. In this work, we systematically investigate transformation and adversarial training based defenses…

Sound · Computer Science 2022-06-08 Guangke Chen , Zhe Zhao , Fu Song , Sen Chen , Lingling Fan , Feng Wang , Jiashui Wang

Recent works have revealed the vulnerability of automatic speech recognition (ASR) models to adversarial examples (AEs), i.e., small perturbations that cause an error in the transcription of the audio signal. Studying audio adversarial…

Sound · Computer Science 2022-03-21 Marie Biolková , Bac Nguyen

In this paper, we propose dictionary attacks against speaker verification - a novel attack vector that aims to match a large fraction of speaker population by chance. We introduce a generic formulation of the attack that can be used with…

Sound · Computer Science 2022-12-13 Mirko Marras , Pawel Korus , Anubhav Jain , Nasir Memon

Speaker attribute perturbation offers a feasible approach to asynchronous voice anonymization by employing adversarially perturbed speech as anonymized output. In order to enhance the identity unlinkability among anonymized utterances from…

Sound · Computer Science 2025-08-22 Liping Chen , Chenyang Guo , Rui Wang , Kong Aik Lee , Zhenhua Ling

We present Malafide, a universal adversarial attack against automatic speaker verification (ASV) spoofing countermeasures (CMs). By introducing convolutional noise using an optimised linear time-invariant filter, Malafide attacks can be…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-14 Michele Panariello , Wanying Ge , Hemlata Tak , Massimiliano Todisco , Nicholas Evans

In realistic environments, speech is usually interfered by various noise and reverberation, which dramatically degrades the performance of automatic speech recognition (ASR) systems. To alleviate this issue, the commonest way is to use a…

Sound · Computer Science 2018-05-04 Bin Liu , Shuai Nie , Yaping Zhang , Dengfeng Ke , Shan Liang , Wenju Liu1

Speaker Identification refers to the process of identifying a person using one's voice from a collection of known speakers. Environmental noise, reverberation and distortion make the task of automatic speaker identification challenging as…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-01 Sabbir Ahmed , Nursadul Mamun , Md Azad Hossain

In the past few years, it has been shown that deep learning systems are highly vulnerable under attacks with adversarial examples. Neural-network-based automatic speech recognition (ASR) systems are no exception. Targeted and untargeted…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-07 Matías Pizarro , Dorothea Kolossa , Asja Fischer

In recent years, significant progress has been made in deep model-based automatic speech recognition (ASR), leading to its widespread deployment in the real world. At the same time, adversarial attacks against deep ASR systems are highly…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-04 Christian Heider Nielsen , Zheng-Hua Tan