中文

CNN-LTE:基于标签树嵌入用于音频场景识别的一类 1-X 池化卷积神经网络

神经与进化计算 2016-08-16 v2 计算机视觉与模式识别 机器学习 多媒体 声音

摘要

我们在本报告中描述了提交至 DCASE 2016 挑战赛的音频场景识别系统。首先,给定场景的标签集,自动构建一棵标签树。该类别分类法随后用于特征提取步骤,其中音频场景实例由标签树嵌入图像表示。最后,在图像特征之上学习针对当前任务定制的不同卷积神经网络,以进行场景识别。我们的系统达到了 81.2% 和 83.3% 的总体识别准确率,并在开发和测试数据上分别以 8.7% 和 6.1% 的绝对提升幅度优于 DCASE 2016 基线。

关键词

引用

@article{arxiv.1607.02303,
  title  = {CNN-LTE: a Class of 1-X Pooling Convolutional Neural Networks on Label Tree Embeddings for Audio Scene Recognition},
  author = {Huy Phan and Lars Hertel and Marco Maass and Philipp Koch and Alfred Mertins},
  journal= {arXiv preprint arXiv:1607.02303},
  year   = {2016}
}

备注

Task1 technical report for the DCASE2016 challenge. arXiv admin note: text overlap with arXiv:1606.07908