English

LHGNN: Local-Higher Order Graph Neural Networks For Audio Classification and Tagging

Sound 2025-01-30 v2 Artificial Intelligence Audio and Speech Processing

Abstract

Transformers have set new benchmarks in audio processing tasks, leveraging self-attention mechanisms to capture complex patterns and dependencies within audio data. However, their focus on pairwise interactions limits their ability to process the higher-order relations essential for identifying distinct audio objects. To address this limitation, this work introduces the Local- Higher Order Graph Neural Network (LHGNN), a graph based model that enhances feature understanding by integrating local neighbourhood information with higher-order data from Fuzzy C-Means clusters, thereby capturing a broader spectrum of audio relationships. Evaluation of the model on three publicly available audio datasets shows that it outperforms Transformer-based models across all benchmarks while operating with substantially fewer parameters. Moreover, LHGNN demonstrates a distinct advantage in scenarios lacking ImageNet pretraining, establishing its effectiveness and efficiency in environments where extensive pretraining data is unavailable.

Keywords

Cite

@article{arxiv.2501.03464,
  title  = {LHGNN: Local-Higher Order Graph Neural Networks For Audio Classification and Tagging},
  author = {Shubhr Singh and Emmanouil Benetos and Huy Phan and Dan Stowell},
  journal= {arXiv preprint arXiv:2501.03464},
  year   = {2025}
}
R2 v1 2026-06-28T20:58:16.058Z