中文

基于生成与预测混合模型的实时丢包 concealment

音频与语音处理 2022-05-13 v1 声音

摘要

随着深度语音增强算法近期展现出远超传统方法抑制噪声、混响与回声的能力,注意力正转向丢包 concealment(PLC)问题。PLC 是一项具有挑战性的任务,因为它不仅涉及实时语音合成,还涉及接收音频与合成 concealment 之间的频繁切换。我们提出一种混合神经网络 PLC 架构,其中缺失语音使用以预测模型为条件的生成模型进行合成。所得算法实现了自然的 concealment,质量超越现有常规 PLC 算法,并在 Interspeech 2022 PLC 挑战赛中排名第二。我们表明该方案不仅适用于未压缩音频,也适用于现代语音编解码器。

关键词

引用

@article{arxiv.2205.05785,
  title  = {Real-Time Packet Loss Concealment With Mixed Generative and Predictive Model},
  author = {Jean-Marc Valin and Ahmed Mustafa and Christopher Montgomery and Timothy B. Terriberry and Michael Klingbeil and Paris Smaragdis and Arvindh Krishnaswamy},
  journal= {arXiv preprint arXiv:2205.05785},
  year   = {2022}
}

备注

Submitted to INTERSPEECH 2022