English

Fusion of Audio and Visual Embeddings for Sound Event Localization and Detection

Audio and Speech Processing 2023-12-15 v1 Sound Image and Video Processing

Abstract

Sound event localization and detection (SELD) combines two subtasks: sound event detection (SED) and direction of arrival (DOA) estimation. SELD is usually tackled as an audio-only problem, but visual information has been recently included. Few audio-visual (AV)-SELD works have been published and most employ vision via face/object bounding boxes, or human pose keypoints. In contrast, we explore the integration of audio and visual feature embeddings extracted with pre-trained deep networks. For the visual modality, we tested ResNet50 and Inflated 3D ConvNet (I3D). Our comparison of AV fusion methods includes the AV-Conformer and Cross-Modal Attentive Fusion (CMAF) model. Our best models outperform the DCASE 2023 Task3 audio-only and AV baselines by a wide margin on the development set of the STARSS23 dataset, making them competitive amongst state-of-the-art results of the AV challenge, without model ensembling, heavy data augmentation, or prediction post-processing. Such techniques and further pre-training could be applied as next steps to improve performance.

Keywords

Cite

@article{arxiv.2312.09034,
  title  = {Fusion of Audio and Visual Embeddings for Sound Event Localization and Detection},
  author = {Davide Berghi and Peipei Wu and Jinzheng Zhao and Wenwu Wang and Philip J. B. Jackson},
  journal= {arXiv preprint arXiv:2312.09034},
  year   = {2023}
}

Comments

ICASSP 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)