中文

联合语义知识蒸馏与遮蔽声学建模:提高语音恢复 intelligibility

密码学与安全 2024-09-17 v1 软件工程

摘要

语音恢复旨在恢复高质量且具可理解性的全频段语音,需考虑多种失真。MaskSR 是一种最近提出的生成模型用于此任务。虽然该模型在质量方面取得了高分,但我们所展示的在可理解性方面可以显著提升。我们通过使用预训练的自监督教师模型对语音编码器组件进行语义表示预测,从而提升语音编码器。随后,对遮蔽语言模型进行条件化,以预测编码目标语音低层频谱细节的声学标记。我们表明,在相同的 MaskSR 模型容量和推理时间下,所提出的模型 MaskSR2 可显著降低词错误率,这是衡量可理解性的典型指标。MaskSR2 还在其他模型中实现了具竞争力的词错误率,同时提供了优异的质量。消融研究显示了各种语义表示的有效性。

关键词

引用

@article{arxiv.2409.09356,
  title  = {Towards Robust Detection of Open Source Software Supply Chain Poisoning Attacks in Industry Environments},
  author = {Xinyi Zheng and Chen Wei and Shenao Wang and Yanjie Zhao and Peiming Gao and Yuanchao Zhang and Kailong Wang and Haoyu Wang},
  journal= {arXiv preprint arXiv:2409.09356},
  year   = {2024}
}

备注

To appear in the 39th IEEE/ACM International Conference on Automated Software Engineering (ASE'24 Industry Showcase), October 27-November 1, 2024, Sacramento, CA, USA