Exploring bat song syllable representations in self-supervised audio encoders
Sound
2024-09-20 v1 Artificial Intelligence
Machine Learning
Audio and Speech Processing
Abstract
How well can deep learning models trained on human-generated sounds distinguish between another species' vocalization types? We analyze the encoding of bat song syllables in several self-supervised audio encoders, and find that models pre-trained on human speech generate the most distinctive representations of different syllable types. These findings form first steps towards the application of cross-species transfer learning in bat bioacoustics, as well as an improved understanding of out-of-distribution signal processing in audio encoder models.
Cite
@article{arxiv.2409.12634,
title = {Exploring bat song syllable representations in self-supervised audio encoders},
author = {Marianne de Heer Kloots and Mirjam Knörnschild},
journal= {arXiv preprint arXiv:2409.12634},
year = {2024}
}
Comments
Presented at VIHAR-2024; see https://vihar-2024.vihar.org/