中文

基于典型相关与深度学习的语音可懂度增强模型用于助听技术

音频与语音处理 2022-02-16 v2 声音

摘要

当前基于深度学习(DL)的噪声环境下语音可懂度增强方法通常训练为最小化干净语音与增强语音特征之间的距离。这些方法常能改善语音质量,但缺乏泛化能力,且可能无法在日常噪声场景中提供所需的可懂度。为应对这些挑战,研究者已探索面向可懂度(I-O)的损失函数来训练 DL 方法以实现鲁棒语音增强(SE)。本文中,我们构建了一种新颖的基于典型相关的 I-O 损失函数,以更有效地训练 DL 算法。具体而言,我们提出一种全卷积 SE 模型,其使用改进的典型相关短时客观可懂度(CC-STOI)度量作为训练代价函数。据我们所知,这是首项将典型相关整合进基于 I-O 的损失函数用于 SE 的工作。对比实验结果表明,在处理未见说话人与噪声时,我们提出的基于 CC-STOI 的 SE 框架在标准客观与主观评价指标上均优于使用常规 STOI 和基于距离损失函数训练的 DL 模型。

关键词

引用

@article{arxiv.2202.04172,
  title  = {A Speech Intelligibility Enhancement Model based on Canonical Correlation and Deep Learning for Hearing-Assistive Technologies},
  author = {Tassadaq Hussain and Muhammad Diyan and Mandar Gogate and Kia Dashtipour and Ahsan Adeel and Yu Tsao and Amir Hussain},
  journal= {arXiv preprint arXiv:2202.04172},
  year   = {2022}
}

备注

We would like to withdraw this article because we have accidentally uploaded the revised version of the same article from another account. The updated version is titled "A Novel Speech Intelligibility Enhancement Model based on Canonical Correlation and Deep Learning" (arXiv:2202.05756)