面向声事件检测的音频转换器有效预训练
音频与语音处理
2024-12-02 v2 声音
摘要
我们提出了一种用于音频频谱转换器的预训练管道,用于帧级声事件检测任务。在常规预训练步骤之上,我们添加了一个精心设计的训练程序,用于AudioSet的帧级注释。这包括平衡抽样器、激进的数据增强和集成知识蒸馏。对于五个转换器,我们在AudioSet帧级预测和帧级声事件检测下游任务上均获得了显著的性能提升,确认了我们管道的有效性。我们发布了 resulting checkpoints,供研究人员直接微调以构建用于声事件检测任务的高性能模型。
引用
@article{arxiv.2409.09546,
title = {Effective Pre-Training of Audio Transformers for Sound Event Detection},
author = {Florian Schmid and Tobias Morocutti and Francesco Foscarin and Jan Schlüter and Paul Primus and Gerhard Widmer},
journal= {arXiv preprint arXiv:2409.09546},
year = {2024}
}
备注
Submitted to ICASSP'25. Source code available: https://github.com/fschmid56/PretrainedSED