English

Privacy-Preserving End-to-End Full-Duplex Speech Dialogue Models

Audio and Speech Processing 2026-03-10 v1 Artificial Intelligence Signal Processing

Abstract

End-to-end full-duplex speech models feed user audio through an always-on LLM backbone, yet the speaker privacy implications of their hidden representations remain unexamined. Following the VoicePrivacy 2024 protocol with a lazy-informed attacker, we show that the hidden states of SALM-Duplex and Moshi leak substantial speaker identity across all transformer layers. Layer-wise and turn-wise analyses reveal that leakage persists across all layers, with SALM-Duplex showing stronger leakage in early layers while Moshi leaks uniformly, and that Linkability rises sharply within the first few turns. We propose two streaming anonymization setups using Stream-Voice-Anon: a waveform-level front-end (Anon-W2W) and a feature-domain replacement (Anon-W2F). Anon-W2F raises EER by over 3.5x relative to the discrete encoder baseline (11.2% to 41.0%), approaching the 50% random-chance ceiling, while Anon-W2W retains 78-93% of baseline sBERT across setups with sub-second response latency (FRL under 0.8 s).

Cite

@article{arxiv.2603.08179,
  title  = {Privacy-Preserving End-to-End Full-Duplex Speech Dialogue Models},
  author = {Nikita Kuzmin and Tao Zhong and Jiajun Deng and Yingke Zhu and Tristan Tsoi and Tianxiang Cao and Simon Lui and Kong Aik Lee and Eng Siong Chng},
  journal= {arXiv preprint arXiv:2603.08179},
  year   = {2026}
}
R2 v1 2026-07-01T11:09:58.964Z