English
Related papers

Related papers: The VoicePrivacy 2024 Challenge Evaluation Plan

200 papers

Speech brain-computer interfaces aim to decipher what a person is trying to say from neural activity alone, restoring communication to people with paralysis who have lost the ability to speak intelligibly. The Brain-to-Text Benchmark '24…

The WildSpoof Challenge aims to advance the use of in-the-wild data in two intertwined speech processing tasks. It consists of two parallel tracks: (1) Text-to-Speech (TTS) synthesis for generating spoofed speech, and (2) Spoofing-robust…

Sound · Computer Science 2025-08-26 Yihan Wu , Jee-weon Jung , Hye-jin Shim , Xin Cheng , Xin Wang

Social media has become a platform for people to stand up and raise their voices against social and criminal acts. Vocalization of such information has allowed the investigation and identification of criminals. However, revealing such…

Computation and Language · Computer Science 2022-11-17 Supriti Vijay , Aman Priyanshu

The Multi-target Challenge aims to assess how well current speech technology is able to determine whether or not a recorded utterance was spoken by one of a large number of blacklisted speakers. It is a form of multi-target speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2019-04-10 Suwon Shon , Najim Dehak , Douglas Reynolds , James Glass

The Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. We recently organized a DNS challenge special session at INTERSPEECH and ICASSP…

The Inspirational and Convincing Audio Generation Challenge 2024 (ICAGC 2024) is part of the ISCSLP 2024 Competitions and Challenges track. While current text-to-speech (TTS) technology can generate high-quality audio, its ability to convey…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-01 Ruibo Fu , Rui Liu , Chunyu Qiang , Yingming Gao , Yi Lu , Shuchen Shi , Tao Wang , Ya Li , Zhengqi Wen , Chen Zhang , Hui Bu , Yukun Liu , Xin Qi , Guanjun Li

Using our voices to access, and interact with, online services raises concerns about the trade-offs between convenience, privacy, and security. The conflict between maintaining privacy and ensuring input authenticity has often been hindered…

Computers and Society · Computer Science 2023-02-28 Ranya Aloufi , Hamed Haddadi , David Boyle

We present a thorough analysis of the findings of the latest iteration of the Singing Voice Conversion Challenge, a scientific event aiming to compare and understand different voice conversion systems in a controlled environment. Compared…

The ICASSP 2026 URGENT Challenge advances the series by focusing on universal speech enhancement (SE) systems that handle diverse distortions, domains, and input conditions. This overview paper details the challenge's motivation, task…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-21 Chenda Li , Wei Wang , Marvin Sach , Wangyou Zhang , Kohei Saijo , Samuele Cornell , Yihui Fu , Zhaoheng Ni , Tim Fingscheidt , Shinji Watanabe , Yanmin Qian

In recent years, the remarkable advancements in deep neural networks have brought tremendous convenience. However, the training process of a highly effective model necessitates a substantial quantity of samples, which brings huge potential…

Sound · Computer Science 2024-09-13 Zhisheng Zhang , Pengyang Huang

We introduce machine unlearning for speech tasks, a novel and underexplored research problem that aims to efficiently and effectively remove the influence of specific data from trained speech models without full retraining. This has…

Machine Learning · Computer Science 2025-06-03 Jiali Cheng , Hadi Amiri

The increasing capabilities of deep neural networks for re-identification, combined with the rise in public surveillance in recent years, pose a substantial threat to individual privacy. Event cameras were initially considered as a…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Katharina Bendig , René Schuster , Nicole Thiemer , Karen Joisten , Didier Stricker

In 2024, we will hold a research paper competition (the third Human Understanding AI Paper Challenge) for the research and development of artificial intelligence technologies to understand human daily life. This document introduces the…

Machine Learning · Computer Science 2024-03-26 Se Won Oh , Hyuntae Jeong , Jeong Mook Lim , Seungeun Chung , Kyoung Ju Noh

Anonymizing textual documents is a highly context-sensitive problem: the appropriate balance between privacy protection and utility preservation varies with the data domain, privacy objectives, and downstream application. However, existing…

Computation and Language · Computer Science 2026-04-21 Gabriel Loiseau , Damien Sileo , Damien Riquet , Maxime Meyer , Marc Tommasi

Audio is a rich sensing modality that is useful for a variety of human activity recognition tasks. However, the ubiquitous nature of smartphones and smart speakers with always-on microphones has led to numerous privacy concerns and a lack…

We propose DarkStream, a streaming speech synthesis model for real-time speaker anonymization. To improve content encoding under strict latency constraints, DarkStream combines a causal waveform encoder, a short lookahead buffer, and…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-08 Waris Quamer , Ricardo Gutierrez-Osuna

The Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. This is the 4th DNS challenge, with the previous editions held at INTERSPEECH 2020,…

Over half of the world's population is bilingual and people often communicate under multilingual scenarios. The Face-Voice Association in Multilingual Environments (FAME) 2026 Challenge, held at ICASSP 2026, focuses on developing methods…

Audio has become an increasingly crucial biometric modality due to its ability to provide an intuitive way for humans to interact with machines. It is currently being used for a range of applications, including person authentication to…

Sound · Computer Science 2023-07-14 Rishabh Ranjan , Mayank Vatsa , Richa Singh

Speech emotion recognition (SER), particularly for naturally expressed emotions, remains a challenging computational task. Key challenges include the inherent subjectivity in emotion annotation and the imbalanced distribution of emotion…

Sound · Computer Science 2025-06-03 Tiantian Feng , Thanathai Lertpetchpun , Dani Byrd , Shrikanth Narayanan