中文

用于音乐分类的多任务自监督预训练

声音 2021-02-08 v1 机器学习 音频与语音处理

摘要

深度学习非常依赖数据,尤其是监督学习需要海量标注数据才能良好工作。机器听觉研究常受限于标注数据不足的问题,因为人工标注获取成本高,且音频标注耗时且不够直观。此外,从标注数据集学到的模型往往嵌入了该特定数据集特有的偏置。因此,无监督学习技术成为解决机器听觉问题的流行方法。特别地,一种利用多种手工音频特征重建的自监督学习技术在应用于语音领域(如情感识别与自动语音识别(ASR))时展现了有前景的结果。在本文中,我们应用自监督与多任务学习方法预训练音乐编码器,并探索多种设计选择,包括编码器架构、结合多任务损失的加权机制,以及前置任务的worker选择。我们研究了这些设计选择如何与各类下游音乐分类任务相互作用。我们发现,在预训练中使用多种音乐特定的worker并配合加权机制以平衡损失,有助于提升并泛化到下游任务。

关键词

引用

@article{arxiv.2102.03229,
  title  = {Multi-Task Self-Supervised Pre-Training for Music Classification},
  author = {Ho-Hsiang Wu and Chieh-Chi Kao and Qingming Tang and Ming Sun and Brian McFee and Juan Pablo Bello and Chao Wang},
  journal= {arXiv preprint arXiv:2102.03229},
  year   = {2021}
}

备注

Copyright 2021 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works