Training Sound Event Detection On A Heterogeneous Dataset
Sound
2020-07-09 v1 Audio and Speech Processing
Signal Processing
Abstract
Training a sound event detection algorithm on a heterogeneous dataset including both recorded and synthetic soundscapes that can have various labeling granularity is a non-trivial task that can lead to systems requiring several technical choices. These technical choices are often passed from one system to another without being questioned. We propose to perform a detailed analysis of DCASE 2020 task 4 sound event detection baseline with regards to several aspects such as the type of data used for training, the parameters of the mean-teacher or the transformations applied while generating the synthetic soundscapes. Some of the parameters that are usually used as default are shown to be sub-optimal.
Cite
@article{arxiv.2007.03931,
title = {Training Sound Event Detection On A Heterogeneous Dataset},
author = {Nicolas Turpault and Romain Serizel},
journal= {arXiv preprint arXiv:2007.03931},
year = {2020}
}