基于共享权重和注意力机制的 ResNet-Conformer 网络用于声音事件定位、检测与距离估计
声音
2025-07-25 v1 音频与语音处理
摘要
本技术报告概述了我们对 DCASE 2024 任务 3A 的方法,专注于声音事件定位与检测(SELD)。SELD 通过估计声音事件的定位与检测,为各种机器认知任务提供有价值的见解,例如环境推断、导航以及其他与声音定位相关的应用。今年的挑战赛采用音频-only(轨道 A)或音视频(轨道 B)输入,对标注的真实声场景录音进行评估。今年的一个显著变化是引入距离估计,评估指标相应调整,以进行全面评估。我们的提交物是针对挑战赛的任务 A,即音频-only 轨道。我们的方法利用对数梅尔频谱图和能量向量,并采用多种数据增强技术。我们提出了基于 EINV2[1] 的网络架构,实现了改进的结果:在开发数据集[2,3]的测试集上,F 分数为 40.2%,角度误差(DOA)为 17.7 度,相对距离误差(RDE)为 0.32。
引用
@article{arxiv.2507.17941,
title = {Resnet-conformer network with shared weights and attention mechanism for sound event localization, detection, and distance estimation},
author = {Quoc Thinh Vo and David Han},
journal= {arXiv preprint arXiv:2507.17941},
year = {2025}
}
备注
This paper has been submitted as a technical report outlining our approach to Task 3A of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2024 and can be found in DCASE2024 technical reports