English

MAGENTA: Magnitude and Geometry-ENhanced Training Approach for Robust Long-Tailed Sound Event Localization and Detection

Audio and Speech Processing 2025-09-22 v1 Sound

Abstract

Deep learning-based Sound Event Localization and Detection (SELD) systems degrade significantly on real-world, long-tailed datasets. Standard regression losses bias learning toward frequent classes, causing rare events to be systematically under-recognized. To address this challenge, we introduce MAGENTA (Magnitude And Geometry-ENhanced Training Approach), a unified loss function that counteracts this bias within a physically interpretable vector space. MAGENTA geometrically decomposes the regression error into radial and angular components, enabling targeted, rarity-aware penalties and strengthened directional modeling. Empirically, MAGENTA substantially improves SELD performance on imbalanced real-world data, providing a principled foundation for a new class of geometry-aware SELD objectives. Code is available at: https://github.com/itsjunwei/MAGENTA_ICASSP

Keywords

Cite

@article{arxiv.2509.15599,
  title  = {MAGENTA: Magnitude and Geometry-ENhanced Training Approach for Robust Long-Tailed Sound Event Localization and Detection},
  author = {Jun-Wei Yeow and Ee-Leng Tan and Santi Peksi and Woon-Seng Gan},
  journal= {arXiv preprint arXiv:2509.15599},
  year   = {2025}
}

Comments

This work has been submitted to IEEE ICASSP 2026 for possible publication