CLAP-S: Support Set Based Adaptation for Downstream Fiber-optic Acoustic Recognition
Abstract
Contrastive Language-Audio Pretraining (CLAP) models have demonstrated unprecedented performance in various acoustic signal recognition tasks. Fiber-optic-based acoustic recognition is one of the most important downstream tasks and plays a significant role in environmental sensing. Adapting CLAP for fiber-optic acoustic recognition has become an active research area. As a non-conventional acoustic sensor, fiber-optic acoustic recognition presents a challenging, domain-specific, low-shot deployment environment with significant domain shifts due to unique frequency response and noise characteristics. To address these challenges, we propose a support-based adaptation method, CLAP-S, which linearly interpolates a CLAP Adapter with the Support Set, leveraging both implicit knowledge through fine-tuning and explicit knowledge retrieved from memory for cross-domain generalization. Experimental results show that our method delivers competitive performance on both laboratory-recorded fiber-optic ESC-50 datasets and a real-world fiber-optic gunshot-firework dataset. Our research also provides valuable insights for other downstream acoustic recognition tasks. The code and gunshot-firework dataset are available at https://github.com/Jingchensun/clap-s.
Keywords
Cite
@article{arxiv.2501.09877,
title = {CLAP-S: Support Set Based Adaptation for Downstream Fiber-optic Acoustic Recognition},
author = {Jingchen Sun and Shaobo Han and Wataru Kohno and Changyou Chen},
journal= {arXiv preprint arXiv:2501.09877},
year = {2025}
}
Comments
Accepted to ICASSP 2025