声景在声谱图中的表现:南亚声音的多标签分类
声音
2026-03-10 v1 多媒体
摘要
环境声音分类是环境监测和文化声景分析中日益重要的领域,尤其在南亚这种声音丰富的环境中更为突出。这些地区 presents a unique challenge as multiple natural, human, and cultural sounds often overlap, straining traditional methods that frequently rely on Mel Frequency Cepstral Coefficients (MFCC)。本研究引入一种新颖的声谱图方法,能够更好地捕捉这些复杂的听觉模式。实现一种卷积神经网络(CNN)架构,用于解决该 SAS-KIIT 数据集上困难的多标签、多类分类问题。为 demonstrate robust and comparability, the approach is also validated using the renowned UrbanSound8K dataset. 结果表明,所提出的声谱图方法显著优于现有的MFCC技术,在两个数据集上实现更高的分类准确率。这一改进为在实际应用中构建更稳健和准确的音频分类系统奠定了基础。
引用
@article{arxiv.2603.08154,
title = {Soundscapes in Spectrograms: Pioneering Multilabel Classification for South Asian Sounds},
author = {Sudip Chakrabarty and Pappu Bishwas and Rajdeep Chatterjee and Tathagata Bandyopadhyay and Digonto Biswas and Bibek Howlader},
journal= {arXiv preprint arXiv:2603.08154},
year = {2026}
}