用于声学场景分类的多时间分辨率卷积神经网络
声音
2018-11-13 v1 多媒体
音频与语音处理
摘要
在本文中,我们提出一种深度神经网络架构,用于声学场景分类任务,该架构利用来自梅尔谱图片段递增时间分辨率的信息。该架构由分离的并行卷积神经网络组成,这些网络为每个输入分辨率学习频谱与时域表示。所选分辨率旨在覆盖场景频谱纹理的细粒度特征及其声学事件分布。所提模型相比性能最佳的单分辨率模型绝对提升 3.56%,相比 DCASE 2017 声学场景分类任务基线绝对提升 12.49%。
引用
@article{arxiv.1811.04419,
title = {Multi-Temporal Resolution Convolutional Neural Networks for Acoustic Scene Classification},
author = {Alexander Schindler and Thomas Lidy and Andreas Rauber},
journal= {arXiv preprint arXiv:1811.04419},
year = {2018}
}
备注
In Proceedings of the Detection and Classification of Acoustic Scenes and Events 2017 Workshop (DCASE2017), November 2017