中文

面向声事件检测的音频转换器有效预训练

音频与语音处理 2024-12-02 v2 声音

摘要

我们提出了一种用于音频频谱转换器的预训练管道,用于帧级声事件检测任务。在常规预训练步骤之上,我们添加了一个精心设计的训练程序,用于AudioSet的帧级注释。这包括平衡抽样器、激进的数据增强和集成知识蒸馏。对于五个转换器,我们在AudioSet帧级预测和帧级声事件检测下游任务上均获得了显著的性能提升,确认了我们管道的有效性。我们发布了 resulting checkpoints,供研究人员直接微调以构建用于声事件检测任务的高性能模型。

关键词

引用

@article{arxiv.2409.09546,
  title  = {Effective Pre-Training of Audio Transformers for Sound Event Detection},
  author = {Florian Schmid and Tobias Morocutti and Francesco Foscarin and Jan Schlüter and Paul Primus and Gerhard Widmer},
  journal= {arXiv preprint arXiv:2409.09546},
  year   = {2024}
}

备注

Submitted to ICASSP'25. Source code available: https://github.com/fschmid56/PretrainedSED