English

On The Effect Of Coding Artifacts On Acoustic Scene Classification

Audio and Speech Processing 2021-12-10 v1 Multimedia Sound Signal Processing

Abstract

Previous DCASE challenges contributed to an increase in the performance of acoustic scene classification systems. State-of-the-art classifiers demand significant processing capabilities and memory which is challenging for resource-constrained mobile or IoT edge devices. Thus, it is more likely to deploy these models on more powerful hardware and classify audio recordings previously uploaded (or streamed) from low-power edge devices. In such scenario, the edge device may apply perceptual audio coding to reduce the transmission data rate. This paper explores the effect of perceptual audio coding on the classification performance using a DCASE 2020 challenge contribution [1]. We found that classification accuracy can degrade by up to 57% compared to classifying original (uncompressed) audio. We further demonstrate how lossy audio compression techniques during model training can improve classification accuracy of compressed audio signals even for audio codecs and codec bitrates not included in the training process.

Keywords

Cite

@article{arxiv.2112.04841,
  title  = {On The Effect Of Coding Artifacts On Acoustic Scene Classification},
  author = {Nagashree K. S. Rao and Nils Peters},
  journal= {arXiv preprint arXiv:2112.04841},
  year   = {2021}
}

Comments

paper presented at the 2021 Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE)

R2 v1 2026-06-24T08:10:32.821Z