Anomalous Sound Detection Meets Noise-Aware Self-Supervised Learning
Abstract
In this paper, we introduce noise-aware self-supervised learning (NA-SSL) models for noise-aware anomalous sound detection (NA-ASD). NA-ASD is an ASD task with two-channel audio recordings, where one microphone is located close to the target machine and the other is located farther away to capture noise. For this task, we simulate two-channel recordings using diverse audio datasets and train NA-SSL models to extract clean SSL representations of the close-microphone signal by using the far-microphone recording dominated by background noise as auxiliary information. The NA-SSL models are then used as frontends in the standard ASD framework. Our experimental evaluation on the DCASE 2026 Challenge Task 2 development dataset demonstrates the effectiveness of the NA-SSL framework across three base SSL models (BEATs, EAT, and Dasheng), both with and without discriminative fine-tuning. Furthermore, the challenge results proved the effectiveness of the proposed approach, where the NA-BEATs system won the challenge by a large margin, achieving an official score of 70.24%, while the second-place system achieved 65.46%.
Cite
@article{arxiv.2608.00447,
title = {Anomalous Sound Detection Meets Noise-Aware Self-Supervised Learning},
author = {Takuya Fujimura and Gordon Wichern and Yoshiki Masuyama and Christoph Boeddeker and Kohei Saijo and Julius Richter and Takahiro Edo and Jonathan Le Roux},
journal= {arXiv preprint arXiv:2608.00447},
year = {2026}
}