English

Learning from Taxonomy: Multi-label Few-Shot Classification for Everyday Sound Recognition

Sound 2022-12-20 v1 Audio and Speech Processing

Abstract

Everyday sound recognition aims to infer types of sound events in audio streams. While many works succeeded in training models with high performance in a fully-supervised manner, they are still restricted to the demand of large quantities of labelled data and the range of predefined classes. To overcome these drawbacks, this work firstly curates a new database named FSD-FS for multi-label few-shot audio classification. It then explores how to incorporate audio taxonomy in few-shot learning. Specifically, this work proposes label-dependent prototypical networks (LaD-protonet) to exploit parent-children relationships between labels. Plus, it applies taxonomy-aware label smoothing techniques to boost model performance. Experiments demonstrate that LaD-protonet outperforms original prototypical networks as well as other state-of-the-art methods. Moreover, its performance can be further boosted when combined with taxonomy-aware label smoothing.

Keywords

Cite

@article{arxiv.2212.08952,
  title  = {Learning from Taxonomy: Multi-label Few-Shot Classification for Everyday Sound Recognition},
  author = {Jinhua Liang and Huy Phan and Emmanouil Benetos},
  journal= {arXiv preprint arXiv:2212.08952},
  year   = {2022}
}

Comments

submitted to ICASSP2023

R2 v1 2026-06-28T07:40:30.167Z