中文

Spatial LibriSpeech:用于空间音频学习的增强数据集

声音 2023-08-21 v1 人工智能 机器学习 音频与语音处理

摘要

我们提出 Spatial LibriSpeech,一个包含超过 650 小时 19 通道音频、一阶 ambisonics(高阶声场)和可选干扰噪声的空间音频数据集。Spatial LibriSpeech 专为机器学习模型训练设计,包含声源位置、说话方向、房间声学及几何的标签。Spatial LibriSpeech 通过将 LibriSpeech 样本在 8k+ 个合成房间中辅以 200k+ 种模拟声学条件增强生成。为展示数据集效用,我们在四项空间音频任务上训练模型,在 3D 声源定位上获得 6.60{\deg} 中位绝对误差,距离上 0.43m,T30 上 90.66ms,DRR 估计上 2.74dB。我们展示相同模型对广泛使用的评测数据集泛化良好,例如在 TUT Sound Events 2018 上 3D 声源定位中位绝对误差 12.43{\deg},在 ACE Challenge 上 T30 估计 157.32ms。

关键词

引用

@article{arxiv.2308.09514,
  title  = {Spatial LibriSpeech: An Augmented Dataset for Spatial Audio Learning},
  author = {Miguel Sarabia and Elena Menyaylenko and Alessandro Toso and Skyler Seto and Zakaria Aldeneh and Shadi Pirhosseinloo and Luca Zappella and Barry-John Theobald and Nicholas Apostoloff and Jonathan Sheaffer},
  journal= {arXiv preprint arXiv:2308.09514},
  year   = {2023}
}