English

A Robust framework for sound event localization and detection on real recordings

Sound 2025-12-30 v1

Abstract

This technical report describes the systems submitted to the DCASE2022 challenge task 3: sound event localization and detection (SELD). The task aims to detect occurrences of sound events and specify their class, furthermore estimate their position. Our system utilizes a ResNet-based model under a proposed robust framework for SELD. To guarantee the generalized performance on the real-world sound scenes, we design the total framework with augmentation techniques, a pipeline of mixing datasets from real-world sound scenes and emulations, and test time augmentation. Augmentation techniques and exploitation of external sound sources enable training diverse samples and keeping the opportunity to train the real-world context enough by maintaining the number of the real recording samples in the batch. In addition, we design a test time augmentation and a clustering-based model ensemble method to aggregate confident predictions. Experimental results show that the model under a proposed framework outperforms the baseline methods and achieves competitive performance in real-world sound recordings.

Keywords

Cite

@article{arxiv.2512.22156,
  title  = {A Robust framework for sound event localization and detection on real recordings},
  author = {Jin Sob Kim and Hyun Joon Park and Wooseok Shin and Sung Won Han},
  journal= {arXiv preprint arXiv:2512.22156},
  year   = {2025}
}

Comments

Technical Report submitted to DCASE 2022 Challenge Task 3 (Winner of the Judge's Award)

R2 v1 2026-07-01T08:41:48.641Z