English
Related papers

Related papers: Improving the Speaker Anonymization Evaluation's R…

200 papers

Sound event detection systems are widely used in various applications such as surveillance and environmental monitoring where data is automatically collected, processed, and sent to a cloud for sound recognition. However, this process may…

Sound · Computer Science 2024-01-04 Shayan Gharib , Minh Tran , Diep Luong , Konstantinos Drossos , Tuomas Virtanen

The human voice conveys unique characteristics of an individual, making voice biometrics a key technology for verifying identities in various industries. Despite the impressive progress of speaker recognition systems in terms of accuracy, a…

Sound · Computer Science 2022-08-24 Gianni Fenu , Giacomo Medda , Mirko Marras , Giacomo Meloni

Statistical methods protecting sensitive information or the identity of the data owner have become critical to ensure privacy of individuals as well as of organizations. This paper investigates anonymization methods based on representation…

Machine Learning · Statistics 2018-02-27 Clément Feutry , Pablo Piantanida , Yoshua Bengio , Pierre Duhamel

Spoken question answering (SQA) is challenging due to complex reasoning on top of the spoken documents. The recent studies have also shown the catastrophic impact of automatic speech recognition (ASR) errors on SQA. Therefore, this work…

Computation and Language · Computer Science 2019-04-18 Chia-Hsuan Lee , Yun-Nung Chen , Hung-Yi Lee

Several recently proposed text-to-speech (TTS) models achieved to generate the speech samples with the human-level quality in the single-speaker and multi-speaker TTS scenarios with a set of pre-defined speakers. However, synthesizing a new…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-23 Byoung Jin Choi , Myeonghun Jeong , Minchan Kim , Sung Hwan Mun , Nam Soo Kim

As automatic speaker verification (ASV) systems are vulnerable to spoofing attacks, they are typically used in conjunction with spoofing countermeasure (CM) systems to improve security. For example, the CM can first determine whether the…

Sound · Computer Science 2022-01-25 Anssi Kanervisto , Ville Hautamäki , Tomi Kinnunen , Junichi Yamagishi

Inspite the emerging importance of Speech Emotion Recognition (SER), the state-of-the-art accuracy is quite low and needs improvement to make commercial applications of SER viable. A key underlying reason for the low accuracy is the…

Sound · Computer Science 2020-03-24 Siddique Latif , Rajib Rana , Sara Khalifa , Raja Jurdak , Julien Epps , Björn W. Schuller

Speaker verification systems often degrade significantly when there is a language mismatch between training and testing data. Being able to improve cross-lingual speaker verification system using unlabeled data can greatly increase the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-03 Wei Xia , Jing Huang , John H. L. Hansen

The use of modern vocoders in an analysis/synthesis pipeline allows us to investigate high-quality voice conversion that can be used for privacy purposes. Here, we propose to transform the speaker embedding and the pitch in order to hide…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-27 Paul-Gauthier Noé , Xiaoxiao Miao , Xin Wang , Junichi Yamagishi , Jean-François Bonastre , Driss Matrouf

Recent studies have shown that large language models (LLMs) can infer private user attributes (e.g., age, location, gender) from user-generated text shared online, enabling rapid and large-scale privacy breaches. Existing…

Cryptography and Security · Computer Science 2026-04-21 Dong Yan , Jian Liang , Ran He , Tieniu Tan

Besides its linguistic content, our speech is rich in biometric information that can be inferred by classifiers. Learning privacy-preserving representations for speech signals enables downstream tasks without sharing unnecessary, private…

Sound · Computer Science 2021-06-18 Dimitrios Stoidis , Andrea Cavallaro

Modern zero-shot text-to-speech (TTS) models offer unprecedented expressivity but also pose serious crime risks, as they can synthesize voices of individuals who never consented. In this context, speaker unlearning aims to prevent the…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-29 Myungjin Lee , Eunji Shin , Jiyoung Lee

The ubiquity of offensive and hateful content on online fora necessitates the need for automatic solutions that detect such content competently across target groups. In this paper we show that text classification models trained on large…

Computation and Language · Computer Science 2021-12-08 Darsh J Shah , Sinong Wang , Han Fang , Hao Ma , Luke Zettlemoyer

Sex conversion in speech involves privacy risks from data collection and often leaves residual sex-specific cues in outputs, even when target speaker references are unavailable. We introduce RASO for Reference-free Adversarial Sex…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-05 Yangyang Qu , Michele Panariello , Massimiliano Todisco , Nicholas Evans

Audio is a rich sensing modality that is useful for a variety of human activity recognition tasks. However, the ubiquitous nature of smartphones and smart speakers with always-on microphones has led to numerous privacy concerns and a lack…

Automatic speaker verification (ASV) is one of the core technologies in biometric identification. With the ubiquitous usage of ASV systems in safety-critical applications, more and more malicious attackers attempt to launch adversarial…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-16 Haibin Wu , Xu Li , Andy T. Liu , Zhiyong Wu , Helen Meng , Hung-yi Lee

The rapid progress in personalized speech generation technology, including personalized text-to-speech (TTS) and voice conversion (VC), poses a challenge in distinguishing between generated and real speech for human listeners, resulting in…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-11 Shihao Chen , Liping Chen , Jie Zhang , KongAik Lee , Zhenhua Ling , Lirong Dai

Machine learning approaches for speech enhancement are becoming increasingly expressive, enabling ever more powerful modifications of input signals. In this paper, we demonstrate that this expressiveness introduces a vulnerability: advanced…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-01 Rostislav Makarov , Lea Schönherr , Timo Gerkmann

Speaker, author, and other biometric identification applications often compare a sample's similarity to a database of templates to determine the identity. Given that data may be noisy and similarity measures can be inaccurate, such a…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-01 Tom Bäckström , Mohammad Hassan Vali , My Nguyen , Silas Rech

Speaker diarization systems segment a conversation recording based on the speakers' identity. Such systems can misclassify the speaker of a portion of audio due to a variety of factors, such as speech pattern variation, background noise,…

Sound · Computer Science 2024-06-26 Anurag Chowdhury , Abhinav Misra , Mark C. Fuhs , Monika Woszczyna
‹ Prev 1 4 5 6 7 8 10 Next ›