English

SP-SEDT: Self-supervised Pre-training for Sound Event Detection Transformer

Sound 2022-04-07 v2 Audio and Speech Processing

Abstract

Recently, an event-based end-to-end model (SEDT) has been proposed for sound event detection (SED) and achieves competitive performance. However, compared with the frame-based model, it requires more training data with temporal annotations to improve the localization ability. Synthetic data is an alternative, but it suffers from a great domain gap with real recordings. Inspired by the great success of UP-DETR in object detection, we propose to self-supervisedly pre-train SEDT (SP-SEDT) by detecting random patches (only cropped along the time axis). Experiments on the DCASE2019 task4 dataset show the proposed SP-SEDT can outperform fine-tuned frame-based model. The ablation study is also conducted to investigate the impact of different loss functions and patch size.

Keywords

Cite

@article{arxiv.2111.15222,
  title  = {SP-SEDT: Self-supervised Pre-training for Sound Event Detection Transformer},
  author = {Zhirong Ye and Xiangdong Wang and Hong Liu and Yueliang Qian and Rui Tao and Long Yan and Kazushige Ouchi},
  journal= {arXiv preprint arXiv:2111.15222},
  year   = {2022}
}

Comments

Submitted to interspeech 2022; added experiments for section 4

R2 v1 2026-06-24T07:57:19.378Z