English

Training Sound Event Detection On A Heterogeneous Dataset

Sound 2020-07-09 v1 Audio and Speech Processing Signal Processing

Abstract

Training a sound event detection algorithm on a heterogeneous dataset including both recorded and synthetic soundscapes that can have various labeling granularity is a non-trivial task that can lead to systems requiring several technical choices. These technical choices are often passed from one system to another without being questioned. We propose to perform a detailed analysis of DCASE 2020 task 4 sound event detection baseline with regards to several aspects such as the type of data used for training, the parameters of the mean-teacher or the transformations applied while generating the synthetic soundscapes. Some of the parameters that are usually used as default are shown to be sub-optimal.

Keywords

Cite

@article{arxiv.2007.03931,
  title  = {Training Sound Event Detection On A Heterogeneous Dataset},
  author = {Nicolas Turpault and Romain Serizel},
  journal= {arXiv preprint arXiv:2007.03931},
  year   = {2020}
}
R2 v1 2026-06-23T16:56:32.217Z