改善低资源条件下的声学场景分类
音频与语音处理
2025-04-29 v2 声音
摘要
声学场景分类(ASC)通过音频信号识别环境。本文探讨 ASC 在低资源条件下的应用,提出一种新型模型 DS-FlexiNet,将 MobileNetV2 中的深度可分离卷积与 ResNet 结构的残差连接相结合,在效率与精度之间取得平衡。为解决硬件限制和设备异构性问题,DS-FlexiNet 采用量化感知训练(QAT)进行模型压缩,并通过自动设备脉冲响应(ADIR)和频率混合风格(Freq-MixStyle)等数据增强方法提高跨设备泛化能力。知识蒸馏(KD)从十二个教师模型中进一步提升在未见设备上的性能。该架构包含自定义残差归一化层以处理设备间的域差异,深度可分离卷积降低计算开销而不牺牲特征表征。实验结果表明,DS-FlexiNet 在资源受限条件下在适应性与性能方面均表现卓越。
关键词
引用
@article{arxiv.2412.20722,
title = {Improving Acoustic Scene Classification in Low-Resource Conditions},
author = {Zhi Chen and Yun-Fei Shao and Yong Ma and Mingsheng Wei and Le Zhang and Wei-Qiang Zhang},
journal= {arXiv preprint arXiv:2412.20722},
year = {2025}
}
备注
Copyright 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component