English

Sound Scene Synthesis at the DCASE 2024 Challenge

Artificial Intelligence 2025-01-16 v1 Sound Audio and Speech Processing

Abstract

This paper presents Task 7 at the DCASE 2024 Challenge: sound scene synthesis. Recent advances in sound synthesis and generative models have enabled the creation of realistic and diverse audio content. We introduce a standardized evaluation framework for comparing different sound scene synthesis systems, incorporating both objective and subjective metrics. The challenge attracted four submissions, which are evaluated using the Fr\'echet Audio Distance (FAD) and human perceptual ratings. Our analysis reveals significant insights into the current capabilities and limitations of sound scene synthesis systems, while also highlighting areas for future improvement in this rapidly evolving field.

Keywords

Cite

@article{arxiv.2501.08587,
  title  = {Sound Scene Synthesis at the DCASE 2024 Challenge},
  author = {Mathieu Lagrange and Junwon Lee and Modan Tailleur and Laurie M. Heller and Keunwoo Choi and Brian McFee and Keisuke Imoto and Yuki Okamoto},
  journal= {arXiv preprint arXiv:2501.08587},
  year   = {2025}
}
R2 v1 2026-06-28T21:06:46.822Z