中文

用于声学场景分类的多时间分辨率卷积神经网络

声音 2018-11-13 v1 多媒体 音频与语音处理

摘要

在本文中,我们提出一种深度神经网络架构,用于声学场景分类任务,该架构利用来自梅尔谱图片段递增时间分辨率的信息。该架构由分离的并行卷积神经网络组成,这些网络为每个输入分辨率学习频谱与时域表示。所选分辨率旨在覆盖场景频谱纹理的细粒度特征及其声学事件分布。所提模型相比性能最佳的单分辨率模型绝对提升 3.56%,相比 DCASE 2017 声学场景分类任务基线绝对提升 12.49%。

关键词

引用

@article{arxiv.1811.04419,
  title  = {Multi-Temporal Resolution Convolutional Neural Networks for Acoustic Scene Classification},
  author = {Alexander Schindler and Thomas Lidy and Andreas Rauber},
  journal= {arXiv preprint arXiv:1811.04419},
  year   = {2018}
}

备注

In Proceedings of the Detection and Classification of Acoustic Scenes and Events 2017 Workshop (DCASE2017), November 2017