中文
相关论文

相关论文: End-to-end streaming model for low-latency speech …

200 篇论文

Speech signals contain a lot of sensitive information, such as the speaker's identity, which raises privacy concerns when speech data get collected. Speaker anonymization aims to transform a speech signal to remove the source speaker's…

声音 · 计算机科学 2023-01-16 Pierre Champion , Denis Jouvet , Anthony Larcher

Real-time voice conversion and speaker anonymization require causal, low-latency synthesis without sacrificing intelligibility or naturalness. Current systems have a core representational mismatch: content is time-varying, while speaker…

音频与语音处理 · 电气工程与系统科学 2026-02-11 Waris Quamer , Mu-Ruei Tseng , Ghady Nasrallah , Ricardo Gutierrez-Osuna

We address the challenge of preserving emotional content in streaming speaker anonymization (SA). Neural audio codec language models trained for audio continuation tend to degrade source emotion: content tokens discard emotional…

音频与语音处理 · 电气工程与系统科学 2026-03-09 Nikita Kuzmin , Kong Aik Lee , Eng Siong Chng

Existing privacy-preserving speech representation learning methods target a single application domain. In this paper, we present a novel framework to anonymize utterance-level speech embeddings generated by pre-trained encoders and show its…

音频与语音处理 · 电气工程与系统科学 2023-10-27 Minh Tran , Mohammad Soleymani

Voice Conversion (VC) aims to modify a speaker's timbre while preserving linguistic content. While recent VC models achieve strong performance, most struggle in real-time streaming scenarios due to high latency, dependence on ASR modules,…

声音 · 计算机科学 2025-10-13 Zhao Guo , Ziqian Ning , Guobin Ma , Lei Xie

Given the speech generation framework that represents the speaker attribute with an embedding vector, asynchronous voice anonymization can be achieved by modifying the speaker embedding derived from the original speech. However, the…

音频与语音处理 · 电气工程与系统科学 2025-10-08 Rui Wang , Liping Chen , Kong Aik Lee , Zhengpeng Zha , Zhenhua Ling

End-to-end transformer-based automatic speech recognition (ASR) systems often capture multiple speech traits in their learned representations that are highly entangled, leading to a lack of interpretability. In this study, we propose the…

音频与语音处理 · 电气工程与系统科学 2024-11-28 Pu Wang , Hugo Van hamme

This paper presents a streaming speaker-attributed automatic speech recognition (SA-ASR) model that can recognize ``who spoke what'' with low latency even when multiple people are speaking simultaneously. Our model is based on token-level…

音频与语音处理 · 电气工程与系统科学 2022-07-18 Naoyuki Kanda , Jian Wu , Yu Wu , Xiong Xiao , Zhong Meng , Xiaofei Wang , Yashesh Gaur , Zhuo Chen , Jinyu Li , Takuya Yoshioka

Anonymisation has the goal of manipulating speech signals in order to degrade the reliability of automatic approaches to speaker recognition, while preserving other aspects of speech, such as those relating to intelligibility and…

音频与语音处理 · 电气工程与系统科学 2021-09-02 Jose Patino , Natalia Tomashenko , Massimiliano Todisco , Andreas Nautsch , Nicholas Evans

Speaker attribute perturbation offers a feasible approach to asynchronous voice anonymization by employing adversarially perturbed speech as anonymized output. In order to enhance the identity unlinkability among anonymized utterances from…

声音 · 计算机科学 2025-08-22 Liping Chen , Chenyang Guo , Rui Wang , Kong Aik Lee , Zhenhua Ling

In speaker anonymization, speech recordings are modified in a way that the identity of the speaker remains hidden. While this technology could help to protect the privacy of individuals around the globe, current research restricts this by…

计算与语言 · 计算机科学 2024-10-08 Sarina Meyer , Florian Lux , Ngoc Thang Vu

Automatic Speech Recognition (ASR) has seen remarkable progress, with models like OpenAI Whisper and NVIDIA Canary achieving state-of-the-art (SOTA) performance in offline transcription. However, these models are not designed for streaming…

计算与语言 · 计算机科学 2026-04-07 Tomer Krichli , Bhiksha Raj , Joseph Keshet

Speaker-independent speech recognition systems trained with data from many users are generally robust against speaker variability and work well for a large population of speakers. However, these systems do not always generalize well for…

音频与语音处理 · 电气工程与系统科学 2019-09-17 Khe Chai Sim , Petr Zadrazil , Françoise Beaufays

In this work, we propose a novel method for modeling numerous speakers, which enables expressing the overall characteristics of speakers in detail like a trained multi-speaker model without additional training on the target speaker's…

声音 · 计算机科学 2024-06-03 Jungil Kong , Junmo Lee , Jeongmin Kim , Beomjeong Kim , Jihoon Park , Dohee Kong , Changheon Lee , Sangjin Kim

Encoder-decoder based sequence-to-sequence models have demonstrated state-of-the-art results in end-to-end automatic speech recognition (ASR). Recently, the transformer architecture, which uses self-attention to model temporal context…

声音 · 计算机科学 2020-07-02 Niko Moritz , Takaaki Hori , Jonathan Le Roux

Speech data on the Internet are proliferating exponentially because of the emergence of social media, and the sharing of such personal data raises obvious security and privacy concerns. One solution to mitigate these concerns involves…

音频与语音处理 · 电气工程与系统科学 2022-11-08 Jixun Yao , Qing Wang , Yi Lei , Pengcheng Guo , Lei Xie , Namin Wang , Jie Liu

Sharing real-world speech utterances is key to the training and deployment of voice-based services. However, it also raises privacy risks as speech contains a wealth of personal data. Speaker anonymization aims to remove speaker information…

Voice anonymisation can be used to help protect speaker privacy when speech data is shared with untrusted others. In most practical applications, while the voice identity should be sanitised, other attributes such as the spoken content…

音频与语音处理 · 电气工程与系统科学 2024-08-09 Michele Panariello , Massimiliano Todisco , Nicholas Evans

In this paper we present a Transformer-Transducer model architecture and a training technique to unify streaming and non-streaming speech recognition models into one model. The model is composed of a stack of transformer layers for audio…

声音 · 计算机科学 2020-10-08 Anshuman Tripathi , Jaeyoung Kim , Qian Zhang , Han Lu , Hasim Sak

We present StreamVC, a streaming voice conversion solution that preserves the content and prosody of any source speech while matching the voice timbre from any target speech. Unlike previous approaches, StreamVC produces the resulting…

音频与语音处理 · 电气工程与系统科学 2024-01-09 Yang Yang , Yury Kartynnik , Yunpeng Li , Jiuqiang Tang , Xing Li , George Sung , Matthias Grundmann