CNN-LTE:基于标签树嵌入用于音频场景识别的一类 1-X 池化卷积神经网络
神经与进化计算
2016-08-16 v2 计算机视觉与模式识别
机器学习
多媒体
声音
摘要
我们在本报告中描述了提交至 DCASE 2016 挑战赛的音频场景识别系统。首先,给定场景的标签集,自动构建一棵标签树。该类别分类法随后用于特征提取步骤,其中音频场景实例由标签树嵌入图像表示。最后,在图像特征之上学习针对当前任务定制的不同卷积神经网络,以进行场景识别。我们的系统达到了 81.2% 和 83.3% 的总体识别准确率,并在开发和测试数据上分别以 8.7% 和 6.1% 的绝对提升幅度优于 DCASE 2016 基线。
引用
@article{arxiv.1607.02303,
title = {CNN-LTE: a Class of 1-X Pooling Convolutional Neural Networks on Label Tree Embeddings for Audio Scene Recognition},
author = {Huy Phan and Lars Hertel and Marco Maass and Philipp Koch and Alfred Mertins},
journal= {arXiv preprint arXiv:1607.02303},
year = {2016}
}
备注
Task1 technical report for the DCASE2016 challenge. arXiv admin note: text overlap with arXiv:1606.07908