中文
相关论文

相关论文: Exploring VQ-VAE with Prosody Parameters for Speak…

200 篇论文

We consider the task of unsupervised extraction of meaningful latent representations of speech by applying autoencoding neural networks to speech waveforms. The goal is to learn a representation able to capture high level semantic content…

机器学习 · 计算机科学 2019-09-12 Jan Chorowski , Ron J. Weiss , Samy Bengio , Aäron van den Oord

We present results and analyses from the third VoicePrivacy Challenge held in 2024, which focuses on advancing voice anonymization technologies. The task was to develop a voice anonymization system for speech data that conceals a speaker's…

State-of-the-art approaches to speaker anonymization typically employ some form of perturbation function to conceal speaker information contained within an x-vector embedding, then resynthesize utterances in the voice of a new…

音频与语音处理 · 电气工程与系统科学 2023-06-06 Michele Panariello , Massimiliano Todisco , Nicholas Evans

In speaker anonymization, speech recordings are modified in a way that the identity of the speaker remains hidden. While this technology could help to protect the privacy of individuals around the globe, current research restricts this by…

计算与语言 · 计算机科学 2024-10-08 Sarina Meyer , Florian Lux , Ngoc Thang Vu

Advances in speech technology now allow unprecedented access to personally identifiable information through speech. To protect such information, the differential privacy field has explored ways to anonymize speech while preserving its…

音频与语音处理 · 电气工程与系统科学 2024-09-06 Zexin Cai , Henry Li Xinyuan , Ashi Garg , Leibny Paola García-Perera , Kevin Duh , Sanjeev Khudanpur , Nicholas Andrews , Matthew Wiesner

In our previous work, we proposed a language-independent speaker anonymization system based on self-supervised learning models. Although the system can anonymize speech data of any language, the anonymization was imperfect, and the speech…

声音 · 计算机科学 2022-03-29 Xiaoxiao Miao , Xin Wang , Erica Cooper , Junichi Yamagishi , Natalia Tomashenko

Prosody modeling is important, but still challenging in expressive voice conversion. As prosody is difficult to model, and other factors, e.g., speaker, environment and content, which are entangled with prosody in speech, should be removed…

音频与语音处理 · 电气工程与系统科学 2022-01-04 Wendong Gan , Bolong Wen , Ying Yan , Haitao Chen , Zhichao Wang , Hongqiang Du , Lei Xie , Kaixuan Guo , Hai Li

Collecting speech data is an important step in training speech recognition systems and other speech-based machine learning models. However, the issue of privacy protection is an increasing concern that must be addressed. The current study…

The growing use of voice user interfaces has led to a surge in the collection and storage of speech data. While data collection allows for the development of efficient tools powering most speech services, it also poses serious privacy…

密码学与安全 · 计算机科学 2024-03-04 Pierre Champion

Recent advancements in Text-to-Speech (TTS) systems have enabled the generation of natural and expressive speech from textual input. Accented TTS aims to enhance user experience by making the synthesized speech more relatable to minority…

音频与语音处理 · 电气工程与系统科学 2024-10-18 Jan Melechovsky , Ambuj Mehrish , Berrak Sisman , Dorien Herremans

Over the last decade, the use of Automatic Speaker Verification (ASV) systems has become increasingly widespread in response to the growing need for secure and efficient identity verification methods. The voice data encompasses a wealth of…

音频与语音处理 · 电气工程与系统科学 2023-07-06 Oubaïda Chouchane , Michele Panariello , Oualid Zari , Ismet Kerenciler , Imen Chihaoui , Massimiliano Todisco , Melek Önen

The recently proposed x-vector based anonymization scheme converts any input voice into that of a random pseudo-speaker. In this paper, we present a flexible pseudo-speaker selection technique as a baseline for the first VoicePrivacy…

音频与语音处理 · 电气工程与系统科学 2020-05-19 Brij Mohan Lal Srivastava , Natalia Tomashenko , Xin Wang , Emmanuel Vincent , Junichi Yamagishi , Mohamed Maouche , Aurélien Bellet , Marc Tommasi

We introduce a novel method to improve the performance of the VoicePrivacy Challenge 2022 baseline B1 variants. Among the known deficiencies of x-vector-based anonymization systems is the insufficient disentangling of the input features. In…

音频与语音处理 · 电气工程与系统科学 2022-11-01 Ünal Ege Gaznepoglu , Anna Leschanowsky , Nils Peters

Voice anonymization systems aim to protect speaker privacy by obscuring vocal traits while preserving the linguistic content relevant for downstream applications. However, because these linguistic cues remain intact, they can be exploited…

音频与语音处理 · 电气工程与系统科学 2026-05-21 Ahmad Aloradi , Ünal Ege Gaznepoglu , Emanuël A. P. Habets , Daniel Tenbrinck

Automatic speaker recognition algorithms typically characterize speech audio using short-term spectral features that encode the physiological and anatomical aspects of speech production. Such algorithms do not fully capitalize on…

声音 · 计算机科学 2021-02-16 Anurag Chowdhury , Arun Ross , Prabu David

Adolescent suicide is a critical global health issue, and speech provides a cost-effective modality for automatic suicide risk detection. Given the vulnerable population, protecting speaker identity is particularly important, as speech…

音频与语音处理 · 电气工程与系统科学 2026-01-26 Ziyun Cui , Sike Jia , Yang Lin , Yinan Duan , Diyang Qu , Runsen Chen , Chao Zhang , Chang Lei , Wen Wu

Anonymisation has the goal of manipulating speech signals in order to degrade the reliability of automatic approaches to speaker recognition, while preserving other aspects of speech, such as those relating to intelligibility and…

音频与语音处理 · 电气工程与系统科学 2021-09-02 Jose Patino , Natalia Tomashenko , Massimiliano Todisco , Andreas Nautsch , Nicholas Evans

Recent years have seen remarkable progress in speech emotion recognition (SER), thanks to advances in deep learning techniques. However, the limited availability of labeled data remains a significant challenge in the field. Self-supervised…

声音 · 计算机科学 2023-04-24 Samir Sadok , Simon Leglaive , Renaud Séguier

Current voice conversion (VC) methods can successfully convert timbre of the audio. As modeling source audio's prosody effectively is a challenging task, there are still limitations of transferring source style to the converted speech. This…

音频与语音处理 · 电气工程与系统科学 2021-06-29 Zhichao Wang , Xinyong Zhou , Fengyu Yang , Tao Li , Hongqiang Du , Lei Xie , Wendong Gan , Haitao Chen , Hai Li

Emotion plays a significant role in speech interaction, conveyed through tone, pitch, and rhythm, enabling the expression of feelings and intentions beyond words to create a more personalized experience. However, most existing speaker…

音频与语音处理 · 电气工程与系统科学 2025-06-02 Jixun Yao , Hexin Liu , Eng Siong Chng , Lei Xie