English
Related papers

Related papers: SAIC: Integration of Speech Anonymization and Iden…

200 papers

Social intelligence is essential for understanding and reasoning about human expressions, intents and interactions. One representative benchmark for its study is Social Intelligence Queries (Social-IQ), a dataset of multiple-choice…

Computation and Language · Computer Science 2023-10-31 Xiao-Yu Guo , Yuan-Fang Li , Gholamreza Haffari

This paper presents a novel framework for multi-talker automatic speech recognition without the need for auxiliary information. Serialized Output Training (SOT), a widely used approach, suffers from recognition errors due to speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-10 Asahi Sakuma , Hiroaki Sato , Ryuga Sugano , Tadashi Kumano , Yoshihiko Kawai , Tetsuji Ogawa

This paper proposes a fully explainable approach to speaker verification (SV), a task that fundamentally relies on individual speaker characteristics. The opaque use of speaker attributes in current SV systems raises concerns of trust.…

Sound · Computer Science 2024-05-31 Xiaoliang Wu , Chau Luu , Peter Bell , Ajitha Rajan

Speaker, author, and other biometric identification applications often compare a sample's similarity to a database of templates to determine the identity. Given that data may be noisy and similarity measures can be inaccurate, such a…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-01 Tom Bäckström , Mohammad Hassan Vali , My Nguyen , Silas Rech

We use the term re-identification to refer to the process of recovering the original speaker's identity from anonymized speech outputs. Speaker de-identification systems aim to reduce the risk of re-identification, but most evaluations…

Sound · Computer Science 2025-09-19 Seungmin Seo , Oleg Aulov , P. Jonathon Phillips

Recognizing human non-speech vocalizations is an important task and has broad applications such as automatic sound transcription and health condition monitoring. However, existing datasets have a relatively small number of vocal sound…

Sound · Computer Science 2022-06-22 Yuan Gong , Jin Yu , James Glass

The success of deep learning-based speaker verification systems is largely attributed to access to large-scale and diverse speaker identity data. However, collecting data from more identities is expensive, challenging, and often limited by…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-27 Tianchi Liu , Ruijie Tao , Qiongqiong Wang , Yidi Jiang , Hardik B. Sailor , Ke Zhang , Jingru Lin , Haizhou Li

Automated speaker identification (SID) is a crucial step for the personalization of a wide range of speech-enabled services. Typical SID systems use a symmetric enrollment-verification framework with a single model to derive embeddings both…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-28 Chenyang Gao , Brecht Desplanques , Chelsea J. -T. Ju , Aman Chadha , Andreas Stolcke

Discrete audio tokens have recently gained considerable attention for their potential to bridge audio and language processing, enabling multimodal language models that can both generate and understand audio. However, preserving key…

The 2021 Speaker Recognition Evaluation (SRE21) was the latest cycle of the ongoing evaluation series conducted by the U.S. National Institute of Standards and Technology (NIST) since 1996. It was the second large-scale multimodal…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-22 Seyed Omid Sadjadi , Craig Greenberg , Elliot Singer , Lisa Mason , Douglas Reynolds

Speaker verification (SV) systems are currently being used to make sensitive decisions like giving access to bank accounts or deciding whether the voice of a suspect coincides with that of the perpetrator of a crime. Ensuring that these…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-18 Mariel Estevez , Luciana Ferrer

We propose an approach for training speaker identification models in a weakly supervised manner. We concentrate on the setting where the training data consists of a set of audio recordings and the speaker annotation is provided only at the…

Sound · Computer Science 2018-06-25 Martin Karu , Tanel Alumäe

Voice anonymization has been developed as a technique for preserving privacy by replacing the speaker's voice in a speech signal with that of a pseudo-speaker, thereby obscuring the original voice attributes from machine recognition and…

Sound · Computer Science 2024-11-13 Rui Wang , Liping Chen , Kong AiK Lee , Zhen-Hua Ling

In speaker verification (SV), the acoustic mismatch between children's and adults' speech leads to suboptimal performance when adult-trained SV systems are applied to children's speaker verification (C-SV). While domain adaptation…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-05 Jiusi Zheng , Vishwas Shetty , Natarajan Balaji Shankar , Abeer Alwan

Data anonymization is often a task carried out by humans. Automating it would reduce the cost and time required to complete this task. This paper presents a pipeline to automate the anonymization of audio data in French. We propose a…

Sound · Computer Science 2022-04-28 Guillaume Baril , Patrick Cardinal , Alessandro Lameiras Koerich

Speaker identification in the household scenario (e.g., for smart speakers) is typically based on only a few enrollment utterances but a much larger set of unlabeled data, suggesting semisupervised learning to improve speaker profiles. We…

Sound · Computer Science 2022-02-22 Long Chen , Venkatesh Ravichandran , Andreas Stolcke

The goal of this paper is to learn robust speaker representation for bilingual speaking scenario. The majority of the world's population speak at least two languages; however, most speaker recognition systems fail to recognise the same…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-08 Kihyun Nam , Youkyum Kim , Jaesung Huh , Hee Soo Heo , Jee-weon Jung , Joon Son Chung

Automatic Speaker Diarization (ASD) is an enabling technology with numerous applications, which deals with recordings of multiple speakers, raising special concerns in terms of privacy. In fact, in remote settings, where recordings are…

Audio and Speech Processing · Electrical Eng. & Systems 2023-04-19 Francisco Teixeira , Alberto Abad , Bhiksha Raj , Isabel Trancoso

In speech technologies, speaker's voice representation is used in many applications such as speech recognition, voice conversion, speech synthesis and, obviously, user authentication. Modern vocal representations of the speaker are based on…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-17 Paul-Gauthier Noé , Mohammad Mohammadamini , Driss Matrouf , Titouan Parcollet , Andreas Nautsch , Jean-François Bonastre

Existing privacy-preserving speech representation learning methods target a single application domain. In this paper, we present a novel framework to anonymize utterance-level speech embeddings generated by pre-trained encoders and show its…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-27 Minh Tran , Mohammad Soleymani