中文

SemiPL:一种面向事件声源定位的半监督方法

计算机视觉与模式识别 2024-05-01 v1 多媒体 声音 音频与语音处理

摘要

近年来,Event Sound Source Localization 在 various 领域得到广泛应用。Recent works typically relying on the contrastive learning framework show impressive performance。然而,所有工作都基于 relatively simple 的 large 数据集。在许多应用中,如 crowd management 和 emergency response services,理解和分析人类行为(人物动作和互动)、声音和混乱事件中的声音至关重要。本文将现有模型应用于更 complex 的数据集,探讨参数对模型的影响,并提出一种半监督改进方法 SemiPL。随着 data quantity 的增加以及 label quality 的影响,self-supervised learning 将成为不可阻止的趋势。实验表明,参数调整会正面影响现有模型。特别是,SSPL 在 Chaotic World 中的 cIoU 提升了 12.2%,AUC 提升了 0.56%。代码可在 https://github.com/ly245422/SSPL 获取。

关键词

引用

@article{arxiv.2404.19615,
  title  = {SemiPL: A Semi-supervised Method for Event Sound Source Localization},
  author = {Yue Li and Baiqiao Yin and Jinfu Liu and Jiajun Wen and Jiaying Lin and Mengyuan Liu},
  journal= {arXiv preprint arXiv:2404.19615},
  year   = {2024}
}