中文

用于说话人提取的空间选择性深度非线性滤波器

音频与语音处理 2023-04-05 v2 机器学习 声音

摘要

在多人同时说话的场景中,信号的空间特征是提取目标信号最显著的特性。本工作中,我们开发了一种可指向任意目标方向的深度联合空间-频谱非线性滤波器。为此,我们提出了一种简单有效的调节机制,基于目标方向设置滤波器循环层的初始状态。我们表明该方案比基线方法更有效,并以零性能代价提升了滤波器的灵活性。所得空间选择性非线性滤波器也可用于任意数量说话人的语音分离,并如本文所示可实现非常精准的多说话人定位。

关键词

引用

@article{arxiv.2211.02420,
  title  = {Spatially Selective Deep Non-linear Filters for Speaker Extraction},
  author = {Kristina Tesch and Timo Gerkmann},
  journal= {arXiv preprint arXiv:2211.02420},
  year   = {2023}
}

备注

\copyright 2023 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works